What this is

A deterministic, single-file response-cache laboratory. Replay the same multi-tenant workload under exact, semantic, or hybrid lookup; tune threshold, TTL, context keys, and fact-change invalidation; then inspect the candidate, score components, key boundary, entry age, answer, and correctness classification behind every hit or miss.

Why this is mind-blowing

The fastest-looking policy can be the least useful one. The observatory makes that failure measurable: stale and contaminated hits lower precision and contribute zero trusted savings, while an immutable tenant guard proves that aggressive reuse never becomes cross-tenant leakage. The “semantic” model is deliberately local and lexical, with every weight and penalty visible instead of hidden behind an imaginary embedding API.