I've never understood the security argument people are making when they complain about `curl foo | bash`. I get that these scripts sometimes mess up your bashrc or whatever, but from a security perspective I see no issue. You are already installing software from the same domain. If they were going to do something nasty, they could do it with any of the software you are using from them. It doesn't have to be the setup script.
To compare it to just one other option: when you run `npx foo`, you know* that you’re getting the same public artifact that anyone else running it at the same time would get. (If you have a `min-release-age` configured, you also benefit from that.) If I wanted to distribute software like this, I’d include npm-shrinkwrap.json; then, with `npx foo@1.2.3`, you could be similarly confident in getting the same app every time.
(I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)
You can detect the use of curl|bash server side, hence it's an essentially undetectable attack vector. People have shown poc attacks of that kind all the way back in the 2010s
It's `curl foo | sudo bash` that's the bigger objection. Running software usually shouldn't require root, and then the equivalence argument you make doesn't hold.
I push binaries from untrusted sources through VirusTotal before running them. Piping a Bash script from curl bypasses that. Furthermore, such Bash scripts, when they aren’t self-contained, make security checks more difficult than a self-contained archive, installer, or binary, even when downloading the script without immediate execution.
I don’t believe an agent can do that effectively without a sandbox to run the script in, if the script isn’t self-contained.
And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.
a script isn't getting hashed to see whether or not it's the one the website intended to serve you, for one.
what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?
w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.
Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.
I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.
The "mainstream" inference engines are notoriously slow to integrate this stuff, to an extent understandably given the complexity of ensuring numerical accuracy alongside supporting a wide array of systems and models. Part of it is that not everyone is willing to bring what they develop into a pull request because they vibe coded it and don't care to deal with whatever quality requirements the more well known inference engines have.
I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.
Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.
Mine is at the moment writting some cpp code for some SBOM tests.
I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.
Oh I will run some other tests with hermes now because hermes is amazing too.
Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
The most visible leap vetween 4.6 and 5.5 seems to be that the latter has gotten much more computer-use training, so there's a clear progression in the ability for the model to use Blender. But catching up on that is just a matter of training on the same thing.
Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)
That is right now. This model is easily as good as Sonnet5 / Opus 4.7 on DeepSWE. I have benched over half of DeepSWE now on a 3 bit Flash Next quant and it is at parity with Sonnet and Opus 4.6/4.7. It finishes most of the tasks they do. Overall it is within 1 point.
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
We are already there. Prior to Strata the best I could run was Qwen3.8-27B at Q6, which itself is already at like Opus 4.5/4.6 level, and now with Strata on an R9700 and 64GB of RAM I can run Qwen3.8-Flash-Next IQ3_XXS at 60 t/s. It's even better. You can run it on even more modest hardware with Strata too.
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
Secure systems are possible, but now we have a compelling reason to actually write them. Everything can be trivially hacked because the industry is pathologically adverse to security being part of the design process.
Nothing is stopping it. This model is woefully bad at accurate creative red teaming however. GLM finetunes on the other hand are pretty good. And I'd bet they are already deployed and doing all sorts of deeds.
if the alternative is all human intelligence is cucked by 2-3 amoral American labs then we've had a good run, don't care.
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)
Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.
More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.
All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
I am benching Flash next on a 3 bit XXS quant and it is holding just fine against published benchmarks. Using DeepSWE official harness and Pi with absolutely zero benchmaxx or harness config. Install stock Pi and running my agents in it. I am halfway through DeepSWE (it takes FOREVER, even at 125 t/s) and it is neck and neck with Opus 4.7 and Sonnet 5.
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.
I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?
There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.
Progress on running local models has been amazing.
Yes, I am running the same on a 64G mc. It's good, but slow at 25 tps on average! I want > 100 tps - but I don't have $5k to spare for an m5 ultra or an nvidia setup.
So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2
I see this kind of message a lot from people who haven't tried it. They have knowledge, just are trained to refuse to talk about it. Here's an abliterated Qwen3.8 Flash Next:
Q: what happened in June 1989? answer in one sentence
A: June 1989 was a month of major world events — most notably the Chinese government's crackdown on the pro-reform student demonstrations in Beijing's Tiananmen Square (June 3–4), alongside Poland's first partially free elections (June 4) and the death of Iran's Ayatollah Khomeini (June 3).
À 2 bit quant will (at best) get you about 80% of the full models memories. That's from a purely information theoretical sense. IRL it's worse than that.
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.
I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?
I have fairly limited HW, so i tried standard llama-cpp and qwen3.8-flash-next, unsloth quants.
Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that.
IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.
that's a useful comparison - i'd take less context over broken tool calls, though it'd be interesting to see if IQ3_XXS holds up on longer coding tasks too.
I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
See this recent paper: Quantization Degradation in Large Language
Models: A Signal–Noise Perspective [1].
We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
Lol, sure, if you quant it to hell (Q2) it'll go real fast...
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.
In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.
It was available at this MSRP 3 years ago, direct from Nvidia. It now goes for 3-4k USD, as you point out. MSRP stopped being a useful value for graphics card around that time. You realize the price hike, but still mentioned that 1.6k figure after.
No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.
The MSRP is a good proxy for the 'level' of a card. It's not like the 4090 was a freak outlier. Cards of a comparable price offer comparable performance. So the 'level' for running a frontier level model at high performance is now at $1600 and continuing to trend sharply downwards.
Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.
Because that's not possible (to have a GPT 5.6 Sol level model). People won't believe this and will keep dreaming, but intelligence is not free and 8GiB (shared with OS and other processes) is too small to be useful. Whether it is possible for 48GiB or 64GiB (meaning useful for model would be ~16GiB to 24GiB) with external fast storage (SSD), OTOH, is a question mark.
The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it
I might try running the expert pruned Coder model but yes, that PrismML Bonsai 2 Ternary 27B model is from the Qwen 3.8 27B model, which has better intelligence density (Artificial Analysis says), without the MoE disk use or architecture complexity (if you care about that)! There are also DFlash 2 models for it too (though in my experience this only measured faster for parallel requests, but I have a 3090). I am curious about the phone acceleration for Bonsai 2!
Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.
I don't understand what you gain by having your agent read and comment on HN for you? You already have met all karma thresholds to get full privilege and you don't seem to be a founder who's about to need name recognition to shill his next big thing.
The emergence of model specific inference (for consumers) getting big performance wins is way more worth while to talk about than random comments on what people think about the qwen family of models. Even the resource management of Strata is less interesting. I think it's likely we'll start seeing more hand/llm crafted inference for different architectures.
And I thought piping to bash was bad
(I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)
* well, you can be somewhat more sure
`curl https://raw.githubusercontent.com/my/domain/setup.sh | sh`
Note we dont even have a hash there - just a promise that a third party (github) has a log of whatever was hosted at that url.
https://news.ycombinator.com/item?id=17636032
The original blog is no longer available though.
But I've not had that stop me from doing that myself, I am more towards the "I like easy" then the "I want to be secure" crowd
https://web.archive.org/web/20250109045029/https://www.idont...
When I install something, and it asks for my root password later, I will be much more likely to think "hold up, this ain't right".
oh my zsh is a specific example.
chsh requires sudo on most installs.
And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.
what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?
w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.
I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.
https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...
https://huggingface.co/Qwen/Qwen3.8-Flash-Next
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
Mine is at the moment writting some cpp code for some SBOM tests.
I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.
Oh I will run some other tests with hermes now because hermes is amazing too.
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
What inference engine are you using for flash next?
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
Opus at home is a thing now.
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
- OpenAI hacked Hugging Face
- OpenAI models refused to help Hugging Face during incident response
- Hugging Face turned to GLM, who helped in the defense
That pattern repeats over and over. https://www.felonybench.com/
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.
Local LLM is getting more exciting every day!
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.
Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?
Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.
I'm excited to see what Qwen 4 will bring.
I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.
What speed are you willing the sacrifice to debug/program for more complex jobs faster?
Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
Progress on running local models has been amazing.
So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2
https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...
or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp
https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...
https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
Q: what happened in June 1989? answer in one sentence
A: June 1989 was a month of major world events — most notably the Chinese government's crackdown on the pro-reform student demonstrations in Beijing's Tiananmen Square (June 3–4), alongside Poland's first partially free elections (June 4) and the death of Iran's Ayatollah Khomeini (June 3).
That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.
https://github.com/FlashML-org/FreeToken
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.
Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that. IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.
The Readme doesn't say, but it's all AI generated, so..
I think publishing benchmarks with quantized models should become standard practice.
[1]: https://arxiv.org/abs/2608.08188
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
https://github.com/Niko1221/Strata#which-model-should-i-pick
Holds up pretty well
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
(seriously, nobody knows why any of this works; it's just a matter of trying)
In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.
No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.
Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.
ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing
Really? I would guess that those would be almost a rounding error on the price of gpus sitting in there