Why RAG will outlive the context window debate
Longer context doesn't make retrieval irrelevant — it changes what retrieval is for.
Arguments on AI systems, engineering decisions, and the edges of what machines can do. Specific positions, not trend summaries.
Longer context doesn't make retrieval irrelevant — it changes what retrieval is for.
When the model can write its own tests, passing tests stops being evidence of correctness.
Every agent decision that can't be reviewed is a gap in your production architecture.
Dense embeddings feel like magic until you need to recall a specific invoice number.
The skill isn't writing better prompts. It's knowing when the prompt is the wrong abstraction entirely.
Treating the context window like a conversation history is the first mistake most teams make when building agents.
JSON mode doesn't save you from an underspecified schema. It just makes the failure more consistent.
Most teams choose the most capable model and then discover their latency requirements. That's backwards.
You can't improve what you can't measure. In LLM systems, measurement is the hard part.