Index A | B | C | D | E | F | G | H | I | K | L | M | N | O | P | Q | R | S | T | V | W | X A A2A activation checkpointing agentic RAG AIBrix ARC-AGI arithmetic intensity attention autoregressive AWQ B best-of-N BF16 BitNet Blackwell Blackwell / GB200 BM25 Bradley-Terry C calibration chunked prefill chunking Claude Sonnet 5 / Claude Fable 5 ColBERT compute capability compute-bound computer use contamination continuous batching corrective RAG cross-modal adapter D DAPO DDP decode dense retrieval disaggregated prefill/decode DoRA DPO draft model DSPy dynamic batching E Elo Expected Calibration Error (ECE) expert parallel F finite-state machine (FSM) FlashAttention FlashAttention-3 FLOPs FP16 FP32 FP4 / NV-FP4 FP8 FSDP G Gemini 3.5 Flash Gemini Spark GPQA GPT-5.6 GPTQ GQA gradient accumulation Grok 4.5 GRPO guardrail H handoff HBM HLE HyDE I inference-time scaling INT2 / ternary INT8 / INT4 K kernel KV cache KV eviction KV quantization L LiveCodeBench Llama 4 LLM-as-judge LMCache LoRA M MCP memory-bound MHA Microsoft Agent Framework MiMo-V2.5 mixed precision MLA MoE N nDCG NF4 NIAH NIXL NVIDIA Dynamo NVLink O ORPO P PagedAttention parallel coordinated reasoning pass@k PegaFlow perplexity pipeline parallel prefill Pydantic AI Q QLoRA QuaRot / SpinQuant Qwen3 R RAG ReAct reasoning model reasoning tokens recall@k reranking ridge intensity RLHF RLVR RMSNorm roofline RoPE RRF RULER S SFT SGLang SigLIP SIMT SLO SM smolagents SmoothQuant sparse retrieval speculative decoding SPLADE static batching structured outputs SWE-bench T target model tensor core tensor parallel Terminal-Bench test-time compute TFLOPs thinking budget thinking tokens ThunderKittens tool use TPOT TTFT V Vera Rubin / Rubin GPU vision encoder VLA vLLM V2 VLM W warp X XGrammar