| Thinking | default ON for think-capable models in plain chat (json/tools excluded); reasoning_content vs content; gpt-oss Harmony channels map the same way. Think budget guarantees answer headroom (started_in_think was a three-bug chain). |
| Thinking vs rendered prompt | pipeline: intent -> Jinja -> reconcile_thinking_with_prompt_tail; the prompt tail is ground truth. A closed <think></think> template block treated as "thinking" puts the answer in reasoning_content. enable_thinking:true honored via force_thinking. Do not re-tighten mentions_thinking() (load-bearing for spontaneous-<think> models). |
reasoning_effort | threaded through ChatTemplate::apply*/render_jinja + server snapshot (#1750); identical prompt-token counts across efforts = it is not reaching the template |
| Regex constraints | response_format: {"type":"regex"} or guided_regex; RegexNfa (src/compute/json_schema.h) + RegexConstrainer; lookaround, anchors, \b, backrefs refused in the constrainer (the NFA parses some of them) |
| GBNF constraints | {"type":"grammar"}, top-level grammar, guided_grammar; pushdown simulator src/compute/gbnf_grammar.{h,cpp} + gbnf_parser.cpp, GrammarConstrainer; left recursion, undefined rules, no root, absurd bounds refused at compile time (400); cold mask = 151k-vocab walk, so the per-state mask cache and interned-stack memo are load-bearing |
| JSON schema | sim_advance in src/compute/schema_constrain.{h,cu} is the single grammar source; $ref/$defs; exact-JSON termination |
| Adding a constrainer | close every bypass: InferenceState sites in engine_decode_pipeline.cpp AND engine_scheduler.cpp, constrained-pipeline state in engine_graph_decode.cpp; the ONE apply_constraint_mask helper in src/exec/executor.cu; spec-ngram and graph-loop eligibility (engine_spec_ngram.cpp, engine_graph_decode.cpp); Engine::ensure_constraints_ and pipeline_row_eligible_; the thinking default in handlers_chat_core.cpp; ConstraintManager is POOLED (active flags + per-constrainer caches must reset; GBNF's memoised transitions once leaked into the next grammar) |
| Tool calling | tool-arg enforcement is FSM-backed since #1002 (forced/required, strict:true, parallel_tool_calls on ChatML-JSON; forced on Llama3); the Qwen-Coder/Qwen3.6 XML dialect (<function=/<parameter= raw-text bodies) gets the XML grammar, never the JSON body FSM (which would mask raw newlines in code args); degen_suite constrained covers forced tool_choice |
| Constrained perf | category prefilter + in-string shortcut; ConstrainedPipeline enqueues forward N+1 before the host FSM advances (~102 -> 235 tok/s). Falls back to eager for logprobs, min_p, typical_p, mirostat, DRY, logit_bias, MTP, batch>1 |
cache_control / prefix cache | pins prompt-KV blocks (server.prefix_pin_budget_pct 25, FIFO); reports cache_read/cache_creation_input_tokens; last marked block bounds the pin; TTL accepted, not modeled; PrefixCacheE2ETest is the gate. Hybrids need server.recurrent_snapshot_mb (256 = 3 slabs of 79.5 MiB on the 27B) + host tier server.recurrent_snapshot_host_mb (2048 = 25 slabs; 8 interleaved sessions x 3 turns: turn-2 TTFT 324 -> 163 ms, #1854) |
| Speculation | speculative.mtp_k=-1 auto (#1809): single-stream (max_batch_size=1) + head + not deterministic -> mtp_k=2, ngram=false; concurrent serving declines (head 0.79 GiB per batch slot); reads the RESOLVED batch size (resolve_max_batch_size(): flag > imp.conf > per-load, #1811). Manual few-stream serving: --set speculative.mtp_k=2 --set speculative.ngram=false as a PAIR. Adaptive depth speculative.mtp_adaptive_k on. n-gram default ON dense, off MoE. imp_spec_drafted_total = 1 after an essay = dead |
| Long context | attention.sparse_topk_tokens (off; 4096-8192 for long-ctx serving, floor 8192 on Qwen3.8 for NIAH 8/10), sparse_min_ctx 12288; the key min/max pool must stay VRAM-resident (a kv_cache.max_blocks pin without headroom spilled every prefill kernel +11%). kv_cache.growable opt-in (#1794: 32x8k+512 wall -24%). --prompt-file for 32k prompts |
| Priority | "priority": int body field (vLLM semantics, lower = earlier, default 0), primary sort key, aging within a class, admission only (#1803). Smoke needs a 400-token occupier or the slot frees before the high-prio request lands |
| Request ids | client X-Request-Id echoed on every response (sanitized, 128 chars); server completion id otherwise; JSONL client_request_id |
| OTLP tracing | server.otlp_endpoint=http://host:4318/v1/traces (off by default), server.otlp_service_name; one SERVER span per generation with queue/prefill/decode children, joined to the caller's traceparent; unsampled (-00) not exported; rejected requests (4xx/429/503) emit no span; no traceresponse header. Test: tests/test_server_tracing.py (in make test-server; its collector binds 4318, use IMP_OTLP_PORT=4319 beside Jaeger). Jaeger recipe: --add-host=host.docker.internal:host-gateway, read back GET :16686/api/traces/<id>; Jaeger storage is in-memory. Compare span attributes against the response usage (that is what found imp.cached_tokens = 0, #1856) |
| Model swapping | server.model_swap on: another model name in the models dir swaps (drain, never cancel; failed load restores the previous). WSL2/WDDM never returns a process's peak VRAM: each swap costs the previous footprint (30927 -> 23113 MiB free after one cycle); restart a long-lived swapping server |
| Streaming UTF-8 | Utf8Stitch holds partial bytes before any consumer; holdback_decision cuts at codepoint boundaries; a new sink must not reintroduce a byte cut |
| TCP | TCP_NODELAY on accepted sockets (#1803; delayed ACK cost up to ~40 ms ITL for network clients) |
| Serving loop | deferred token delivery (#1758; push_token does not wake the SSE handler inline; diagnose with diagnostics.step_timing, #1759, whose sample phase includes the GPU sync), ragged prefill batching (runtime.prefill_batch, src/runtime/engine_prefill_ragged.cpp, #1780; serial for vision, constraints, logprobs, embeddings, rerank; off for Mamba2, MLA, MTP), prefill pacing (256-token floor charged once per ragged forward, #1781; runtime.prefill_chunk_decode_cap 1024, 4096 for bursts = +~10%), id-based rotor (#1762), graph prewarm (runtime.graph_prewarm, #1761, ~2.3 s at init), scheduler TUs split in #1782 (engine_prefill.cpp, engine_prefill_ragged.cpp, engine_decode_pipeline.cpp), hybrid decode pipeline and prefill |
| Stop handling | the server stops on turn markers at high temperature; keep the guard |
Config keys (struct Server) | server.prefix_cache, prefix_pin_budget_pct, green_contexts (off, sm_120 race), model_swap, model_swap_drain_ms, recurrent_snapshot_mb, recurrent_snapshot_host_mb, otlp_endpoint, otlp_service_name |