All
Articles 128,708Blog Posts 133,462Tech Tutorials 33,276Research Papers 25,143News 18,235
⚡ AI Lessons
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
5d ago
Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
Below Upstream Status sections are from https://github.com/PrismML-Eng/Bonsai-demo Upstream Status for Binary Q1_0 is supported out of the box in upstream llama

Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
5d ago
ggml-zendnn : add Q8_0 quantization support by z-sachin · Pull Request #23414 · ggml-org/llama.cpp
<img src="https://external-preview.redd.it/ywbUHer5UPqXllCBpufhsYRs22uW8BD6WRrxanrHoo4.png?width=640&crop=smart&auto=webp&s=8439de7957b36bbbb323a058
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
5d ago
OvisOCR2 (0.8B): first end-to-end model to top OmniDocBench - I threw 827 real scanned medical docs at it, here's everything I learned
What it is: ATH-MaaS/OvisOCR2 - a 0.8B document-parsing VLM post-trained from Qwen3.5-0.8B (SFT + RL + OPD), Apache 2.0, runs on vLLM 0.22.1. One prompt per pag
Reddit r/LocalLLaMA
🤖 AI Agents & Automation
⚡ AI Lesson
1w ago
GLM-5.2 fearmongering in the press
I don't know where this is headed, but I don't like it. https://futurism.com/artificial-intelligence/open-source-ai-model-scary-mythos GLM-5.2 can be downloaded
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
Need help building a rig / estimating performance for big LLMs to run when fully offline
Hello. So, first, ill probably explain the context, and why we (well, I kinda have a team, so ill say we) want to build something like that. We live in Russia.
Reddit r/LocalLLaMA
📣 Digital Marketing & Growth
⚡ AI Lesson
1w ago
Now brothers we know why we are so fucked up
Samsung chip division's single-year profits beat its past 40 years of profits, combined, due to increased memory and storage prices — Samsung passes Nvidia to b
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
[audio.cpp] What Does the Fox Say: 4 ASR models (Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR) in native C++/GGML, init streaming support, and 327s of audio transcribed in 2.17s.
I just pushed a new audio.cpp update with streaming support and 4 ASR/STT models: Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR (da only). Ov

Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
Which open models help the eco system more?
<a href="https://artificialanalysis.ai/evaluations/artificial-analysis-openness-in
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
Qwen3.5 122B is the best?
I’m using Opencode and a computer with 128gb. So maybe the results would be different on system. I’ve exhaustingly tried Qwen3.6 27B and Qwen3.6 33B. I have no
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
What GUI-first coding tool tool are you pairing your local LLMs with? Opencode isn't it for me.
I've grown very frustrated with OpenCode. The web GUI and desktop app ideas are good, but the execution not so much. The GUI is lacking so many basic features.
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
TL;DR: Full GLM-5.2 (753B MoE) quantized to Int4-Int8Mix + NVFP4 4-bit KV cache, TP=4 across 4× DGX Spark (GB10) at 100K context , run on Terminal-Bench 2.1 wit

Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
Complete local model asset generation pipeline
So I figured I'd update the community given I just shipped a nice littl
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
AI has completely revolutionized how I play RPGs
Crossposting this here because I thought you guys might appreciate it. When ChatGPT and other open source LLMs first came out, there was a lot of speculation as
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
GLM-5.2-504B-Code-GGUF is an imatrix calibrated quant of the most liked REAP, which is 0xSero/GLM-5.2-REAP-504B-GGUF . I also uploaded the imatrix file and the

Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
HF Viewer tons of new features!
I'm glad to announce that we have completely revamped the site

Reddit r/LocalLLaMA
📣 Digital Marketing & Growth
⚡ AI Lesson
1w ago
Seasonic PSU calculator now mentions RTX 5080 SUPER (24GB), RTX 5070 Ti SUPER (24GB) and RTX 5070 SUPER (18GB)
<img src="https://external-preview.redd.it/EXXS4FJQeWr23ez3UDBmR78VMcsO6_H0PWUHicwnC6c.jpeg?width=640&crop=smart&auto=webp&s=fbd9f98a9229c4e9873d997
Reddit r/LocalLLaMA
💻 AI-Assisted Coding
⚡ AI Lesson
1w ago
Qwen3.6-27b does not understand software architechure.
Been using this for real software development for a commercial app. i.e. Not a single file HTML app. I mean a large scale 100k+ loc project that needs proper ar
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
QLLM, no transformer, no mamba and new noval architecture with O(1) inference is finally out as model
okay so you might be following me or not.. but I have been working in AI since last 10+ years and our first product in AI was released in 2014 https://web.archi

Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
Döner Bench round 2: Quant compare
I reiterated on the <a href="https://www.reddit.com/r

Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
Can you trust local models to answer accurately?
My goal is to improve as a developer, thus I needed to know if loca
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
1w ago
China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model
https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model According to The Information, MiniMax plans to launc
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
2w ago
Best Local VLMs - July 2026
Share what your favorite models are right now and why . Given the nature of the beast in evaluating VLMs (untrustworthiness of benchmarks, immature tooling, int
Reddit r/LocalLLaMA
🔍 RAG & Vector Search
⚡ AI Lesson
2w ago
What's in your RAG?
I want to up my game, with RAG. I tested it ages ago, but haven't found a usecase. I play with coding, projects, and light sysadmin work. # Thoughts RFC library
Reddit r/LocalLLaMA
🧠 Large Language Models
⚡ AI Lesson
2w ago
My reasons to run local models
I can finetune any model on any dataset I want. I can use techniques like speculative decoding and other sota approaches to get the max tps The llm provides lik
DeepCamp AI