close

DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Embeddings Cannot Say No: An Intent Detector's Real Numbers

Embeddings Cannot Say No: An Intent Detector's Real Numbers

1
Comments
5 min read
How LangChain extracted value out of their own agents

How LangChain extracted value out of their own agents

Comments
1 min read
Node.js API Key Text Classification: JSON Validation Before Multi-Provider Gateway Failover

Node.js API Key Text Classification: JSON Validation Before Multi-Provider Gateway Failover

Comments
6 min read
Adding Responses API to an Agent Framework

Adding Responses API to an Agent Framework

Comments
6 min read
A judge that agrees with your humans 92 percent of the time can be at 60 percent where the gate actually decides

A judge that agrees with your humans 92 percent of the time can be at 60 percent where the gate actually decides

Comments
8 min read
The Model Reading My Benchmark Mattered More Than the Memory System Did

The Model Reading My Benchmark Mattered More Than the Memory System Did

Comments
7 min read
Guard Implementation Patterns to Stop AI Agent Runaway Behavior — 7 Types Extracted from Real-World Logs

Guard Implementation Patterns to Stop AI Agent Runaway Behavior — 7 Types Extracted from Real-World Logs

Comments
5 min read
Testing OmniRoute locally: verify your AI gateway actually works before you trust it

Testing OmniRoute locally: verify your AI gateway actually works before you trust it

Comments
4 min read
Did FP8 make the model dumber? A per-prompt regression check for quantized serving

Did FP8 make the model dumber? A per-prompt regression check for quantized serving

Comments
3 min read
A Better FP4 Gradient Quantizer That Training Couldn't Notice

A Better FP4 Gradient Quantizer That Training Couldn't Notice

Comments
7 min read
Claude Opus 5 Is Too Good to Stay in a Browser—Put It in Discord, Slack, Telegram & LINE

Claude Opus 5 Is Too Good to Stay in a Browser—Put It in Discord, Slack, Telegram & LINE

1
Comments
3 min read
Needle 2: the 14 MB agentic model, tested properly

Needle 2: the 14 MB agentic model, tested properly

1
Comments
1 min read
Computer use leaves beta, request shape changes — the weekly AI engineering brief

Computer use leaves beta, request shape changes — the weekly AI engineering brief

Comments
6 min read
How I Built a Reliable LLM Pipeline for Ad Creative Evaluation (with Strict Pydantic Contracts)

How I Built a Reliable LLM Pipeline for Ad Creative Evaluation (with Strict Pydantic Contracts)

1
Comments 1
3 min read
Passing Once Isn't Reliable — This Week's Agent Engineering Puts the Harness Before the Model

Passing Once Isn't Reliable — This Week's Agent Engineering Puts the Harness Before the Model

Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.