

I guess botched remasters aside lol, there isn’t really a limit to how poorly something can be done


I guess botched remasters aside lol, there isn’t really a limit to how poorly something can be done


There are other ways investigators could have connected activity to Ono. He may have used Nyaa without a VPN while logged into an account. Data from website cookies or content-delivery-network logs could also have been matched with Nyaa activity and tracker records. CODA has not said whether any of those methods were used.
VPN isn’t enough, you gotta worry about fingerprinting and cookies too. Recently it was revealed that Windows 11 has strong fingerprinting options that could be used to track you.
I wonder if they were using a VPN or seedbox. You could even do remote desktop into seedbox and browse the sites from there.


Old music doesn’t get re-recorded for digital releases, they pull it from the original master again, which is better and more authentic than vinyl.
I mean if you went to a concert would you want them to add hiss and pop to the performance so it sounds more like vinyl? Lol certainly you wouldn’t want them to intentionally worsen the signal-to-noise ratio
my laptop is crappy, so like 5 tokens per second lol, prompt processing of like 20 tokens per second
I think a decent laptop nowadays, even running CPU only, could probably do like 5x faster
I’ve run Qwen 3.5 4b and Gemma 4 e2b on CPU only, this should be faster than those I think (fewer active parameters). If you have AVX512 or AVX10 then it should help a bit. Still slow compared to a GPU lol.
anyone try this? this might be good for my crappy laptop lol
is it good enough to use with Zoo Code? is it better than Qwen 3.5 4b?
EDIT: woa

But not yet supported in llama.cpp https://github.com/ggml-org/llama.cpp/pull/26608


Actually funny he’s not asking it to work harder (that would be system prompt or user message), he’s forcing it to think that it will work harder


That’s a really cool idea. It’s like inception for an LLM, you make it think it was the one that thought of this lol


Have you tried preserve thinking? https://lemmus.org/post/24365786


In a few minutes a significant performance improvement incoming
👀
this has been a crazy few weeks! lol


true, it’s not perfectly clear
also I just saw this



have you tried Qwen 3.6 35b a3b? check my guide, it’s still relevant to you just with different numbers because you have 12GB


Gemma is probably good for that, as long as it’s consistently succeeding at the tool calls.


make sure that holds up with large context, you might need to step down to Q3 (which I’ve heard is still good for this model, many people are even using IQ2)


You’re looking for “2160p remux” torrents


this appears to be the untouched model in GGUF format, 180 GB, MXFP4
https://huggingface.co/bartowski/DeepSeek-V4-Flash-0731-GGUF


the original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher


--n-cpu-moe 36 --spec-type draft-mtp --spec-draft-n-max 3 does seem to speed up token generation for me
Can’t use llama-bench for MTP. In a basic tests it seems to improve from about 26 to 30 tokens per second output. But it seems to hurt my input speed from about 1300 pp down to 800.
https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF
just hoping for a new 35b a3b