* if fully automated agentic delivery was pushed by C-levels to save eng cost - they must be responsible
* if code was written with AI assisted tools, then engineer is responsible
* if engineering manager and PMs pushed hard to release the feature with cutting too much scope, then they should be responsible.
BUT we don't live in ideal world, so here is what would happen (timeline):
* incident started, and getting too costly
* CTO, then eventually CEO joins the incident meeting and starts teaching people how to handle incidents
* incident will cost some money (or multiple of them), post mortem will contain non-sense to hide issues, because they can't blame C-level for pushing their shiny automated JIRA ticket closer agent
* CEO announces layoffs, says sorry, takes all the responsibility for this issue, but kicks off 40% of the team
Yeah sure, have you been in incidents costing millions of dollars because CEO pushed something to deliver fast and same CEO sitting on incident meeting and shouting to everyone why can't they fix this dumb incident?
No but I’ve been in post mortems where ceo’s have acknowledged that a culture of reliability was needed and then they followed through with a subsequent reset of expectations around process.
This proves solving open problems in Mathematics were a search problem.
But it might be not good for human brains, because we trained our brains with these problems and our brain optimized search space in some ways, and yes, we also couldn't solve some these problems.
Now imagine someone gets stuck with a problem which could become its own theory, but they will solve it with LLMs and move on to solve their primary problem, because they don't realize how other problem was a big deal. If theory is not formalized, then it won't contribute to the search space for other person, solving different problem.
All in all:
* people's brain will be shaped differently
* we will lose search space optimizations in our brains
* we lose new theory contributions, which increases the search space to help solve other problems
Kimi K3 is more like "weights available" in that you can download and use them but it is under a custom license that has a bunch of limitations where you have to pay Moonshot for doing some stuff. GLM 5.2 on the other hand is plain old MIT.
Not sure how Qwen3.8-Max is going to be licensed, hopefully it'll be Apache like the smaller ones.
You can do whatever you want with the model within your own organization. If you use it commercially—either as a model-as-a-service business or in a very large-scale product—you should check the additional license terms, which go beyond MIT. My interpretation is that Moonshot cares about the exact inference behavior and accurate representation of their model or derivatives, and perhaps also about capturing some additional value despite their own GPU limitations, so the extra license terms focus on those large-scale commercial deployments.
Considering that very few orgs are going to be able to host a 3T parameter model internally, chances are most deployments would be subject to these restrictions and require a separate license from Moonshot.
US AI companies are already sweating and 100% pressuring the Trump administration for more anti-Chinese regulation, since there have already been talk of Trump considering banning Chinese models. There's however another push back from the startup industry urging them not to ban it, since it will stifle the innovation. In other recent news OpenAI also greatly cut their model prices, 20% for 5.6 Terra and 80% for 5.6 Luna, to stay competitive.
> In other recent news OpenAI also greatly cut their model prices, 20% for 5.6 Terra and 80% for 5.6 Luna, to stay competitive.
I’ve seen comments on HN saying how bad this is for the Chinese model developers since the cheaper option like Deepseek Flash are not longer as price competitive to justify the hassle/risk/lack of multimodal… but isn’t this a gigantic red flag for OpenAI/Anthropic at their current valuations?
Sure, it’s just the lowest end for now, and the enterprise money is at the top of the market. And there’s protectionism/enterprise lock-in/etc that complicate things somewhat.
But still, if the US AI labs ever tap the training brakes for a millisecond, the “inference is still a money maker” argument seems to evaporate when they’ll immediately have to fight a race to the bottom until margins are virtually nothing.
Or if the benchmaxing “line goes up” FOMO mindset starts to lose its luster and companies find their individual niches for productive use of AI and stop bothering with all the latest and greatest churn for top dollar.
Which might be even worse if it means the training arms race is still ongoing but neither Anthropic or OpenAI want to be the first to “lose”. While the marginal value of each new model training run keeps decreasing and enterprises signal they’re more concerned with cost reductions than solving ARC-AGI-7 puzzles.
Huh? V4Flash is still incredibly price competitive. The new version is right around GLM 5.2 and maybe slightly worse than opus 4.8 depending on which benchmark you use while being much cheaper(even factoring how most US zdr providers charge 10x Deepseek’s api caching price). K3 is also a tad behind fable/sol while being alot cheaper
I dont think they have a hard time resisting they did it a while ago, my company mandated everyone delete anything Chinese or Chinese derived back in April I think for no reason than "unsafe"
I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models.
It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.
I get where you're coming from, and the intent to make it easier for people to find examples and verses, but there's a fine line with LLMs giving you answers, is that it's interpreting it in some form. Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?
I don't necessarily mean reguritating it, but choosing which part of the scripture to surface to the user is already some interpretation/choice. Even the devil can quote scripture (I'm playing the devil's advocate here).
Very true. The verses selected come from a tool call. The LLM queries the app for relevant verses. It picks the keywords - so perhaps there’s bias there but the verses are handed to the LLM.
But there are ways to control and constrain the LLMs and what the user is presented with.
These are all top of mind for me and why I felt there could be a better option than asking ChatGPT directly.
> Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?
I think it depends, Catholics wouldn't be able to use this because the Magisterium is the ultimate authority on interpreting Scripture, so the personal interpretation isn't really needed. This is not to say that Catholics don't read the Bible, they are encouraged to do so since it deepens their faith
On the other hand, for Protestant it varies, the High Church denominations are closer to Catholics (though none of them accept the Magisterium) in terms of scripture interpretation, but the Low Church ones (like Baptists or Non-Denominational ) are more open to personal interpenetration.
Disclaimer: I'm a Catholic, so if I made a mistake here fellow Protestants, please correct me.
Is the difference between this and a frontier model that the scripture is guaranteed to be real?
I'm on a team that develops a Bible study app, and we're all relatively content with how the basic models converse regarding scripture. Even as far back as GPT-4 was excellent. They occasionally have minor hallucinations (a dealbreaker for a production app), but they do an excellent job with theology and Bible scholarship, given reasonable guardrails.
I'll admit I'm coming from the perspective of "should we be implementing this?" It seems, on the surface, that a strong embedding-based verse retrieval covers the bases at a microfraction of the cost.
If you're interested, check out the development server where we're working on this. You navigate to the search (magnifying glass) and then hit "Meaning". Sorry for the confusing route; we're still deciding on back-end details and haven't focused on the front yet.
I tried using Deepseek with Claude Code and it under performed. The use I was referring to was within a product and I feel it's a lot of value for the cost.
If that's marketed well it feels like it should cause a system shock like R1 did. It would also be interesting to see the reaction with code models becoming so good already i.e. cost efficient models aren't necessarily invalidated early by progress that matters.
Open flash model is competing against OpenAI's 'Sonnet' model at the price of GPT 3, I am really excited about this release, hopefully it holds up in real work as well
IIRC GPT 3 was priced at per 1k tokens, had to check, the biggest GPT 3 model from OpenAI was $0.06/1k, so $60 / 1M. gpt-3.5-turbo was the first model after ChatGPT and that was $2 / 1M. And no caching. So not really in the same ballpark
these are impressive findings, I am curious what was your process to convert existing code to formal verification languages like TLA+.
My basic understanding was to verify high level abstractions (e.g. transport ACK, fsyncs and so on), but verifying this deep probably requires complete verification of stdlib methods used by Postgres, otherwise how can you pinpoint culprit is the sscanf?
Right now I'm only doing very small simple functions. Kani[0] takes care of translating the code to an intermediate representation for me. It converts the Rust code and C code into a GOTO program[1] which verifiers can then run on top of
GOTO is back ! So glad to see CBMC used. I used to write translators to GOTO for simple code checking and was wondering where the recent state of the art was. Thanks for the pointers.
Did you have a look at why3 and generating verification conditions from Rust or C code (as frama-c does) ?
Imagine you pass the interview and should work with these people.
1. Your life belongs to them
2. You must prioritize company over everything (according to their wishes), for exchange of 0.000001% of company shares which can get diluted anytime
3. New Chief Marketing decides to rebrand and your tattoo and part of your body belongs to trash history (who knows maybe they will require you to get new tattoo)
* if fully automated agentic delivery was pushed by C-levels to save eng cost - they must be responsible
* if code was written with AI assisted tools, then engineer is responsible
* if engineering manager and PMs pushed hard to release the feature with cutting too much scope, then they should be responsible.
BUT we don't live in ideal world, so here is what would happen (timeline):
* incident started, and getting too costly
* CTO, then eventually CEO joins the incident meeting and starts teaching people how to handle incidents
* incident will cost some money (or multiple of them), post mortem will contain non-sense to hide issues, because they can't blame C-level for pushing their shiny automated JIRA ticket closer agent
* CEO announces layoffs, says sorry, takes all the responsibility for this issue, but kicks off 40% of the team
* We are hiring...
reply