All
Articles 130,411Blog Posts 135,035Tech Tutorials 33,729Research Papers 25,422News 18,473
⚡ AI Lessons

Dev.to · Ethan Walker
🧠 Large Language Models
⚡ AI Lesson
1w ago
Your LLM-as-judge disagrees with itself between runs
Same outputs, same judge, two runs, two scores. The gate flickered red then green on a branch with...

Dev.to · Ethan Walker
🧠 Large Language Models
⚡ AI Lesson
1w ago
LLM-as-judge disagrees with itself between runs
The flap I had a faithfulness gate on merge: judge scores every case, the mean has to clear 0.80....

Dev.to · Ethan Walker
🧠 Large Language Models
⚡ AI Lesson
2w ago
When an LLM answer is wrong, the trace is where you look. Some tools make that easy.
A user reports a hallucinated answer in prod. To fix it you need the full trace of that one request,...

Dev.to · Ethan Walker
☁️ DevOps & Cloud
⚡ AI Lesson
2w ago
our CI passed. Your agent isn't operator-ready.
Your CI passed. Your agent isn't operator-ready. We shipped a document-extraction agent to...

Dev.to · Ethan Walker
📐 ML Fundamentals
⚡ AI Lesson
4w ago
91% pass rate. Gate green. Shipped. Worst regression we had all quarter.
The gate was a fixed 90% threshold on an intent-classification eval. The change came in at 91%,...

Dev.to · Ethan Walker
⚡ AI Lesson
1mo ago
Promptfoo is a CI gate, not an eval framework. Treating it like one cost us $4,200
Last Monday I logged into our billing dashboard and saw a $4,200 LangSmith spike from the weekend....
DeepCamp AI