close

DEV Community

jidonglab profile picture

jidonglab

1 project a week. Building and sharing the entire process — from idea to shipped product in 7 days. Currently: AI news automation.

Vibe Coding Is Fine. Vibe Debugging Is What Kills You

Vibe Coding Is Fine. Vibe Debugging Is What Kills You

1
Comments 3
6 min read

Want to connect with jidonglab?

Create an account to connect with jidonglab. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
Why Strong Engineers Fail Coding Interviews: A Scorecard Autopsy

Why Strong Engineers Fail Coding Interviews: A Scorecard Autopsy

Comments
6 min read
Why AI Coding Agents Never Delete Code (I Measured It)

Why AI Coding Agents Never Delete Code (I Measured It)

Comments
6 min read
Any Questions for Us? 7 Questions to Ask the Interviewer

Any Questions for Us? 7 Questions to Ask the Interviewer

Comments
6 min read
Using AI in a Live Coding Interview: How Interviewers Know

Using AI in a Live Coding Interview: How Interviewers Know

Comments
6 min read
Gradient Accumulation Loss Bug: Why accum=8 Isn't Batch Size 8

Gradient Accumulation Loss Bug: Why accum=8 Isn't Batch Size 8

Comments
7 min read
Filtered Vector Search: Why HNSW Recall Collapses at 1% Selectivity

Filtered Vector Search: Why HNSW Recall Collapses at 1% Selectivity

Comments
7 min read
YaRN vs NTK RoPE Scaling: Why 4x Context Breaks Short Prompts

YaRN vs NTK RoPE Scaling: Why 4x Context Breaks Short Prompts

Comments
7 min read
Attention Sinks: Why Sliding-Window KV Eviction Breaks Your LLM

Attention Sinks: Why Sliding-Window KV Eviction Breaks Your LLM

Comments
8 min read
Claude Prompt Caching: Why Agent Loops Miss the 20-Block Lookback

Claude Prompt Caching: Why Agent Loops Miss the 20-Block Lookback

Comments
7 min read
Why temperature=0 Isn't Deterministic: LLM Batch Invariance

Why temperature=0 Isn't Deterministic: LLM Batch Invariance

Comments
8 min read
acc vs acc_norm: Why Length Bias Skews LLM Eval Scores

acc vs acc_norm: Why Length Bias Skews LLM Eval Scores

Comments 1
7 min read
Speculative Decoding: Why Batching Kills Your 3x Speedup

Speculative Decoding: Why Batching Kills Your 3x Speedup

Comments
7 min read
DPO Likelihood Displacement: Why Chosen Logprobs Fall Too

DPO Likelihood Displacement: Why Chosen Logprobs Fall Too

Comments
7 min read
SFT Loss Masking: Why Your Fine-Tuned LLM Never Emits EOS

SFT Loss Masking: Why Your Fine-Tuned LLM Never Emits EOS

Comments
7 min read
LLM Logprob Calibration: Why RLHF Makes Confidence Useless

LLM Logprob Calibration: Why RLHF Makes Confidence Useless

Comments
8 min read
Activation Outliers: Why W8A8 INT8 Quantization Wrecks Accuracy

Activation Outliers: Why W8A8 INT8 Quantization Wrecks Accuracy

Comments
8 min read
Left Padding vs Right Padding: Why Batched LLM Generation Breaks

Left Padding vs Right Padding: Why Batched LLM Generation Breaks

Comments
7 min read
GRPO Zero-Variance Groups: Why Half Your Rollouts Do Nothing

GRPO Zero-Variance Groups: Why Half Your Rollouts Do Nothing

Comments
8 min read
FP8 KV Cache Quantization: Why Keys Break Before Values

FP8 KV Cache Quantization: Why Keys Break Before Values

Comments
7 min read
Sampler Order: Why temperature and top_p Don't Compose

Sampler Order: Why temperature and top_p Don't Compose

Comments
7 min read
Why resize_token_embeddings Breaks Your LLM Fine-Tune

Why resize_token_embeddings Breaks Your LLM Fine-Tune

Comments
7 min read
Reciprocal Rank Fusion: Why k=60 Buries Your Best BM25 Hit

Reciprocal Rank Fusion: Why k=60 Buries Your Best BM25 Hit

Comments
7 min read
Embedding Hubness: Why One Chunk Shows Up in Every RAG Query

Embedding Hubness: Why One Chunk Shows Up in Every RAG Query

Comments
7 min read
LLM-as-Judge Position Bias: Why Swapping A and B Flips Wins

LLM-as-Judge Position Bias: Why Swapping A and B Flips Wins

Comments
7 min read
GQA KV Cache Replication: Why TP=16 Doubles Your Memory

GQA KV Cache Replication: Why TP=16 Doubles Your Memory

Comments
7 min read
MoE Expert Load Imbalance: Why Decode Stalls on One GPU

MoE Expert Load Imbalance: Why Decode Stalls on One GPU

Comments
8 min read
Chunked Prefill: Why max_num_batched_tokens=512 Kills Throughput

Chunked Prefill: Why max_num_batched_tokens=512 Kills Throughput

Comments
7 min read
rsLoRA: Why LoRA alpha/r Scaling Kills High-Rank Fine-Tunes

rsLoRA: Why LoRA alpha/r Scaling Kills High-Rank Fine-Tunes

Comments
7 min read
Sequence Packing: Why Cross-Document Attention Ruins Your SFT

Sequence Packing: Why Cross-Document Attention Ruins Your SFT

Comments
7 min read
BM25 Length Normalization: Why Long RAG Chunks Never Rank

BM25 Length Normalization: Why Long RAG Chunks Never Rank

Comments
7 min read
Matryoshka Embedding Truncation: Why 256 Dims Breaks Thresholds

Matryoshka Embedding Truncation: Why 256 Dims Breaks Thresholds

Comments
7 min read
Attention Entropy: Why Softmax Blurs Retrieval at 128k Context

Attention Entropy: Why Softmax Blurs Retrieval at 128k Context

Comments
7 min read
Digit Tokenization: Why a Comma Changes Your LLM's Math Answer

Digit Tokenization: Why a Comma Changes Your LLM's Math Answer

Comments
6 min read
Token Healing: Why a Trailing Space Breaks Your LLM Output

Token Healing: Why a Trailing Space Breaks Your LLM Output

Comments
7 min read
Grammar-Constrained Decoding: Why Valid JSON Gets Wrong Answers

Grammar-Constrained Decoding: Why Valid JSON Gets Wrong Answers

Comments
7 min read
KV Cache Preemption: Why vLLM Throughput Collapses Under Load

KV Cache Preemption: Why vLLM Throughput Collapses Under Load

Comments
7 min read
Filtered Vector Search: Why Metadata Filters Break HNSW Recall

Filtered Vector Search: Why Metadata Filters Break HNSW Recall

Comments
7 min read
Why Repetition Penalty Breaks JSON and Code Generation

Why Repetition Penalty Breaks JSON and Code Generation

Comments
7 min read
RoPE Scaling: Why Raising rope_theta Breaks Short Context

RoPE Scaling: Why Raising rope_theta Breaks Short Context

Comments
7 min read
Claude Prompt Caching: Why cache_read_input_tokens Stays 0

Claude Prompt Caching: Why cache_read_input_tokens Stays 0

Comments
7 min read
Attention Sinks: Why Evicting Token 0 Wrecks Sliding-Window KV

Attention Sinks: Why Evicting Token 0 Wrecks Sliding-Window KV

Comments
7 min read
Why temperature=0 Still Gives Different Answers: Batch Invariance

Why temperature=0 Still Gives Different Answers: Batch Invariance

Comments
7 min read
Speculative Decoding: Why 80% Acceptance Still Loses at Batch 64

Speculative Decoding: Why 80% Acceptance Still Loses at Batch 64

Comments
8 min read
Why JSON Schema Field Order Breaks Structured Output Accuracy

Why JSON Schema Field Order Breaks Structured Output Accuracy

Comments
7 min read
Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

Comments 1
7 min read
DPO Likelihood Displacement: Why Chosen Responses Get Rarer

DPO Likelihood Displacement: Why Chosen Responses Get Rarer

Comments
7 min read
ColBERT Late Interaction: Why MaxSim Beats One Vector Per Chunk

ColBERT Late Interaction: Why MaxSim Beats One Vector Per Chunk

Comments
7 min read
Surface Form Competition: Why Log-Prob Answer Scoring Fails

Surface Form Competition: Why Log-Prob Answer Scoring Fails

Comments
7 min read
Activation Outliers: Why W8A8 INT8 Quantization Needs SmoothQuant

Activation Outliers: Why W8A8 INT8 Quantization Needs SmoothQuant

Comments
8 min read
Cross-Encoder Reranker Score Calibration: Why 0.5 Cutoffs Fail

Cross-Encoder Reranker Score Calibration: Why 0.5 Cutoffs Fail

Comments
8 min read
MCP Tool Sprawl: Why 40 Tools Wreck Tool Selection Accuracy

MCP Tool Sprawl: Why 40 Tools Wreck Tool Selection Accuracy

Comments
7 min read
Why I measure interview silence with WebAudio, not the Speech API

Why I measure interview silence with WebAudio, not the Speech API

1
Comments
5 min read
Best-of-N Sampling: Why N=64 Scores Higher and Answers Worse

Best-of-N Sampling: Why N=64 Scores Higher and Answers Worse

1
Comments
7 min read
Building an AI mock interviewer that tells you why you'd fail: what the turn engine taught me

Building an AI mock interviewer that tells you why you'd fail: what the turn engine taught me

Comments
5 min read
Reciprocal Rank Fusion: Why k=60 Buries Your Best Hit

Reciprocal Rank Fusion: Why k=60 Buries Your Best Hit

Comments
7 min read
Hard Negative Mining Breaks Embedding Fine-Tunes Without Denoising

Hard Negative Mining Breaks Embedding Fine-Tunes Without Denoising

Comments
7 min read
LLM Eval Noise: Why Your 2-Point Accuracy Win Isn't Real

LLM Eval Noise: Why Your 2-Point Accuracy Win Isn't Real

Comments
7 min read
Grouped Query Attention: Why 8 KV Heads Decide Your Batch Size

Grouped Query Attention: Why 8 KV Heads Decide Your Batch Size

Comments
7 min read
Binary Quantized Embeddings: 32x Smaller, If You Center First

Binary Quantized Embeddings: 32x Smaller, If You Center First

Comments
8 min read
loading...