Trelis Research

Live experiments and interactive results.

BPBench: which LLMs best compress fresh human text

An LLM is a compressor — the better it models language, the fewer bits it needs to encode text it's never seen. BPBench measures that as exact bits-per-byte on fresh, provably-human writing published after each model's training cutoff. The best compressors are open models from Chinese labs; the closed frontier is hard to benchmark at all — and that's a result in itself.

July 2026 · 12 models · contamination-proof prose

LLM Valuation Forecasts by Knowledge Cutoff

Models with knowledge cutoffs from Sep 2021 to Feb 2026 forecast valuations of OpenAI, Anthropic, NVIDIA, Alphabet and Meta — versus what actually happened. Plus: why point estimates mislead, hindsight calibration, and what tail shapes LLMs actually believe in.

July 2026 · 26 models · 2,500+ forecasts