Live experiments and interactive results.
An LLM is a compressor — the better it models language, the fewer bits it needs to encode text it's never seen. BPBench measures that as exact bits-per-byte on fresh, provably-human writing published after each model's training cutoff. The best compressors are open models from Chinese labs; the closed frontier is hard to benchmark at all — and that's a result in itself.
Models with knowledge cutoffs from Sep 2021 to Feb 2026 forecast valuations of OpenAI, Anthropic, NVIDIA, Alphabet and Meta — versus what actually happened. Plus: why point estimates mislead, hindsight calibration, and what tail shapes LLMs actually believe in.