close

DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
LLM evals are a parameter sweep — use a parameter sweep tool

LLM evals are a parameter sweep — use a parameter sweep tool

Comments
9 min read
How to Build a Good Human-in-the-Loop for AI-Driven Deployments

How to Build a Good Human-in-the-Loop for AI-Driven Deployments

1
Comments
7 min read
I Reviewed 12 Free-Tier Integrations. The Same Six Myths Kept Appearing.

I Reviewed 12 Free-Tier Integrations. The Same Six Myths Kept Appearing.

Comments
4 min read
Free Tokens, Real Queues: Measure What Your LLM Calls Actually Cost

Free Tokens, Real Queues: Measure What Your LLM Calls Actually Cost

Comments
4 min read
LLM Evaluation for Software Engineers Without an ML Background: the 90+ Checks We Actually Run

LLM Evaluation for Software Engineers Without an ML Background: the 90+ Checks We Actually Run

Comments
10 min read
Flash Onyx 2.1, one day later: my model spent 400 tokens thinking and returned an empty string

Flash Onyx 2.1, one day later: my model spent 400 tokens thinking and returned an empty string

Comments 1
6 min read
GLM-5.3-Flash: Z.ai Reveals Ox Alpha Was Its Open Multimodal Model

GLM-5.3-Flash: Z.ai Reveals Ox Alpha Was Its Open Multimodal Model

Comments
7 min read
SGLang outputs endless repetition on NVFP4 models: the FP8 lm_head bug

SGLang outputs endless repetition on NVFP4 models: the FP8 lm_head bug

Comments
3 min read
A Decision Tree for Free-Tier AI Automation: Terms, Branches, Worked Leaves

A Decision Tree for Free-Tier AI Automation: Terms, Branches, Worked Leaves

Comments
5 min read
INTRODUCTION TO RAG (RETRIEVAL AUGMENTED GENERATION)

INTRODUCTION TO RAG (RETRIEVAL AUGMENTED GENERATION)

Comments
4 min read
Inside IBM Granite 4.2: Building an Orchestrate-Ready Reasoning App

Inside IBM Granite 4.2: Building an Orchestrate-Ready Reasoning App

Comments
17 min read
DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

Comments
2 min read
Why My LLM Agent Fabricated Numbers From Stale Context

Why My LLM Agent Fabricated Numbers From Stale Context

Comments
6 min read
cached_tokens is 0 because your system prompt isn't stable

cached_tokens is 0 because your system prompt isn't stable

Comments
3 min read
I Seeded Bugs Into My Own PR to Test the AI Reviewer

I Seeded Bugs Into My Own PR to Test the AI Reviewer

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.