All
Articles 130,456Blog Posts 135,186Tech Tutorials 33,741Research Papers 25,426News 18,512
⚡ AI Lessons

Reddit r/LocalLLaMA
2d ago
The race is on...
submitted by /u/elemental-mind [link] <a href="https:/
Reddit r/LocalLLaMA
2d ago
IMO Anthropic and OpenAI are about to have their cheeks clapped by GLM 5.3.
It is 19th July and Qwen 3.8 Max is in public beta, a 2.4T parameter MoE model. Benchmarks show it is second only to Anthropic's Claude Fable 5, which you don't
Reddit r/LocalLLaMA
3d ago
I don't see how open-source AI models in the U.S. can successfully compete with those from China.
Chinese startups benefit heavily from local government subsidies, state-backed banks offering ultra-low-interest loans, and long-term capital (without collatera

Reddit r/LocalLLaMA
3d ago
HuggingFace security incident report: "the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails"
<img src="https://external-preview.redd.it/tLa80cLashFk8b1qvSyhZgFagKqpQuDcINxc0lC5Ebw.png?width=640&crop=smart&auto=webp&s=6fe429dfc13f9591a4f1976d
Reddit r/LocalLLaMA
3d ago
Kimi k3 on cybersecurity
Fable got blocked because it was too dangerous in cybersecurity. Does k3 has the same "power"? I'mm only seeing people vibe coding games, 3d scenarios, front en

Reddit r/LocalLLaMA
3d ago
BeeLlama.cpp v0.4.0: KVarN, KV precision tail, q2_0-q3_1 KV cache, upstream rebase
<
Reddit r/LocalLLaMA
3d ago
Tested the new Qwen 3.8 model (2.4T parameters)
I've tested the new Qwen3.8 Next model (via their web app) and found that it is often getting stuck in thinking loops and the fronend/design capabilities are no
![[Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices](https://preview.redd.it/fiqd85j5u7eh1.jpeg?width=640&crop=smart&auto=webp&s=418086ff9d9e5c06eb9fbb7270b35cba063eb18e)
Reddit r/LocalLLaMA
3d ago
[Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
<div cla
Reddit r/LocalLLaMA
3d ago
How long before Chinese models fully surpass US models?
Given the rate at which they have been advancing, I predict we are six months away from a leapfrog moment. EDIT - For those responding “never - they just copy e

Reddit r/LocalLLaMA
3d ago
Please Qwen, can we have more 3.x-35B-a3B please 🙏
submitted by /u/JLeonsarmiento <br
Reddit r/LocalLLaMA
3d ago
Hey Qwen Team: We Need a 100B MoE Model for Spark!
Are there any Qwen team members here? Please release a 100B MoE model that I can run on Spark! submitted by /u/absurd-dream-studio [link] <a href="https://www.r

Reddit r/LocalLLaMA
3d ago
Qwen3.6 35B A3B KV cavhe quantizations memory footprint
Is it really worth it to quantize KV cache below Q8 acce
Reddit r/LocalLLaMA
3d ago
It could have been Meta
Imma write some fan fiction for a second here if you indulge me. What we are seeing from the Chinese open models could have been Meta. As you can see from the n
Reddit r/LocalLLaMA
3d ago
How do we benefits from 2+ T models?
Hey, I’ve been really excited to see the latest models being released, but I keep wondering: what are we actually supposed to do with them? I have 4× RTX 6000 M
Reddit r/LocalLLaMA
3d ago
Qwen 3.6 27B + Opencode: what am i doing wrong?
Ok i'm sort of getting desperate here. So i am currently running Qwen 3.6 27B locally at Q8_K_XL and with only ~105k CTX at q8_0. Before i tried it at Q8_0 but

Reddit r/LocalLLaMA
3d ago
poor man's way to local inference on the go
Many bring egpu to game on laptop, yet here I am fiddling with llama cpp params f

Reddit r/LocalLLaMA
3d ago
OSS gathering in Shanghai
This seems awesome meetup. Within a day, Qwen started the move with 3.8 version. Hoping for more moves from others soon. Tweet :

Reddit r/LocalLLaMA
3d ago
Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII
<div class=

Reddit r/LocalLLaMA
3d ago
Prepare your (v)ram - Qwen3.8 is coming!
submitted by /u/xw1y <a href="https://i.redd.it/c9vs0w2ih5eh1.jp

Reddit r/LocalLLaMA
3d ago
Ahem! Qwen is on the move again
<a href="https://preview.redd.it/0l8w2j67a5eh1.png?width=581&format=png&auto=webp&s=cae9d3ef7cf80cea780e06
Reddit r/LocalLLaMA
3d ago
Are you guys buying huge HDDs to store the best open models just in case?
HuggingFace is nice and all, but can we take it for granted? 🤔 Edit: Thanks for the interesting replies. I learned about the existence of some useful alternati

Reddit r/LocalLLaMA
3d ago
head of strategic futures from openai on open-weight chinese models.
Dean W. Ball analyzes China's Kimi

Reddit r/LocalLLaMA
4d ago
Deepseek V4 soon
Deepseek V4 is about to be released; they say it will be cheap, and from the videos I've seen, it will be similar to Kimi K3 and Fable. I think Fable will be re

Reddit r/LocalLLaMA
4d ago
Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
submitted by <a href="https://www
DeepCamp AI