From 100,000 Documents to Grounded Answers
A production-oriented RAG architecture: tiered parsing, canonical document IR, metadata and ACL filtering, BM25 plus HNSW, RRF, reranking, evaluation, and selective GraphRAG.
Field notes from building AI systems
Practical writing on retrieval-augmented generation, reinforcement learning, applied GenAI, and the engineering decisions behind scalable machine learning systems.
Architecture first, vendor choices second.
Production RAG · Long-form guide
A practical design for ingesting mixed PDFs, scans, images, tables, and diagrams—then retrieving them with metadata-aware BM25 and semantic search, reciprocal rank fusion, reranking, and selective GraphRAG.
Search by title, summary, or topic.
A production-oriented RAG architecture: tiered parsing, canonical document IR, metadata and ACL filtering, BM25 plus HNSW, RRF, reranking, evaluation, and selective GraphRAG.
Play, pause, and scrub through the entire production RAG pipeline—from raw mixed documents to a streamed, cited answer.
The art of designing systems that grow gracefully instead of collapsing under their own success.
Why a system handling 100,000 requests per second can still deliver a terrible user experience—and why performance tuning matters at every layer.
Try another search or choose a different topic.