how to ingest live web data for LLMs, reduce RAG hallucination, manage crawling latency, and stream structured JSON context directly to AI models.