MangoCats

joined 2 years ago
[–] MangoCats@feddit.it 2 points 1 month ago

I could believe that it stays warm enough that ice can’t stay on the surface

When the sun shines, yes. After 30 days of straight overcast? Not so much.

[–] MangoCats@feddit.it 4 points 1 month ago

Got news about snow and ice in Switzerland: it happens in the valleys during winter just as heavy as it ever does at the higher elevations.

[–] MangoCats@feddit.it 6 points 1 month ago

Not just on the dam itself, but across the lake to reduce evaporation - like they've been doing extensively in Australia.

[–] MangoCats@feddit.it 3 points 1 month ago

I mean, the dam is already wired into the power grid, the top of the dam gets far more hours of sun than the valleys, it's almost as if "someone" didn't think about things before being amazed at the outcome.

[–] MangoCats@feddit.it 3 points 1 month ago

Putting solar panels in a valley in Switzerland is... a graphic demonstration of tunnel vision.

[–] MangoCats@feddit.it 2 points 1 month ago

I tried your link, got a Grok 4.2 v Grok 4.3 pairing, they did O.K. 4.2 better than 4.3 at making a webserver that queries and re-serves NEXRAD data, but... they didn't get too far down the feature list (history of rainfall graphs) before they fell apart.

[–] MangoCats@feddit.it 1 points 1 month ago

Yeah, like the rest of LLM output, I'd say 80% of an average code review is worth considering - and if that includes anything you might have otherwise missed, that's a win compared to learning about the problem post-launch.

[–] MangoCats@feddit.it 0 points 1 month ago

We're getting to a point where code failures are causing deaths, large numbers of deaths... Boeing's MAX10 was a code / design / training / management failure.

[–] MangoCats@feddit.it 0 points 1 month ago (3 children)

And the stuff it is good at are not bottlenecks.

Disagree. Code review done right is a virtually impossible bottleneck for most companies to handle, so in the past they didn't do it and we have the shitshow of security and other bugs that we are experiencing today.

If you removed LLMs from the face of the planet today, code quality will not suffer significantly.

But LLMs aren't going anywhere, and they are already being used to find vulnerabilities by hats both white and black. They are also finding functional bugs that affect life safety and financial stability.

Code was fine for four decades

The way that municipal water was fine for centuries before chlorination. Cholera outbreaks were just one of those unavoidable realities, like mass school shooting events in the US.

if a machine learning model was purpose made to do code review, instead of general purpose LLMs doing it, how much better it could be, if we actually leveraged what they’re good at.

From what I have seen over the past year, this type of specialization is happening, and the progress is real and significant.

[–] MangoCats@feddit.it 2 points 1 month ago (2 children)

What kind of challenges are you giving arena.ai? When I tried Grok (latest at the time available in Cursor) vs Opus 4.5, Grok was much faster to respond, and hillariously terrible at wrting functional code - but confident in its responses.

[–] MangoCats@feddit.it 1 points 1 month ago

Grok for images is pretty competitive, in the little use I've put it to.

Grok for code is comically bad.

Of course, in the little I have used LLMs for stable diffusion / Control Net image generation, they all suck, Grok is just near the top of the steaming pile. Meanwhile, the best of the code writing and reviewing models are starting to become super-human, in the way that AlphaGo became super-human a few years back at playing Go.

Caveat: writing and reviewing code is just one small part of Software Engineering, the way that pigment mixing is one small part of painting.

[–] MangoCats@feddit.it 1 points 1 month ago

A lot has been made about the tokens being subsidized, and at the frontier models I believe that's very true. I also believe that we're getting to where the frontier may be better, but not necessary, in order to get value out of using the LLM - and a couple of steps back from the frontier is becoming quite affordable now.

I'd be very interested to try out a GLM-5.2 instance on about a $100K server (8 bit quantitized) - which should cost under $20K per year to operate and maintain (although with component prices inflating the way they are that number continues to climb)... if that setup is as useful as, say, Claude Opus 4.5, that's a reasonable price level for a system that should be able to serve a department of 6-10 users pretty well. If GLM-5.2 isn't "there yet" - it's likely just a matter of time before one gets there.

view more: ‹ prev next ›