Anthropic claims that Mythos has low low low hallucination rate of just 40% in its System Card.
LLMs outputting code solving this exact task is a compound function of luck, with non determined a priori chance of success and unknown a priori cost.
This is my main disappointment with agentic coding. Prompting is fine, quality assurance is there from the start. Greenfield, couldn't care less. Established enterprise code? This is a minefield.
If the tokens were 10 to 100 times cheaper, then it would be a maybe.
I also hate how it makes half of my senior engineers dumber.
? Never said that.