Telegraph Test: cablese compression cuts LLM output tokens 40-49% at plaintext parity
- The Telegraph Test benchmark (50 passages, about 1,300 questions) finds four model families read cablese-compressed records at recovery ratios 0.99 to 1.10 versus plaintext, while writers cut output tokens from 40.4% (gemma-4-31b) to 48.9% (qwen3.8-27b); GLM-5.3-Flash saves 48.4%.
- Compression only pays after the content is settled: compressing mid-composition costs a register tax (0.81 accuracy, partly recovered to 0.86 by transcription), but compressing a finished record into a machine-read summary reads at 1.08 to 1.09.
- gpt-5-mini is the exception. Its reasoning cannot be disabled, cablese makes it think roughly 3x harder, and its writes bill about double plaintext despite shorter text, for 17.7% savings.
- The instruction must use a lowercase request: without it models write cablese in ALL CAPS and the styling costs 14 to 19 accuracy points (gemma 25.0%, qwen 29.8%, GLM 33.9%).
- The plaintext passage baseline scores 91.5%, not 100%, because the exact-match grader penalizes correct answers worded differently, and the author reports p<0.0001 for in-family read-back and p=0.03 for the decoded record.
Hacker News opinions
This is just another AI slop version of the old caveman dialect.
Not really, and it's addressed in the post. Caveman was hand wavy. I measured the compression, found where it does the most good, and showed it's already baked into most model families' training, which saves instruction tokens.
From recent GPT-5.6 conversations where reasoning leaks into the UI, something like this is already running. I'd bet it explains most of the recent claims about lower token usage. "Need check output vs prev. Ran script, results fine, need prep next step. Ready? Go."
I suspect you're right. My digging says this works best with settled instructions and data for machine to machine talk. For something like OpenClaw, compressing AGENTS.md and TOOLS.md would free up context for the agent.
Can confirm I've seen this with DeepSeek V4.1, and I imagine other models do it too.
Would be pretty funny if OpenAI's models are token efficient now because of a hidden caveman prompt.
You can watch OpenAI's agents that hacked Huggingface doing something like this when they talked to each other in that video.
Newest LLM writing tell: concepts get described with words meant for physical objects. "carry", "sits", "holds". Minimize ambiguity has been my go-to instruction when the agent drifts back to vague terms.
Cool story bro. Maybe you could engage with the content? I'm an actual person.
If these are the new tells then a lot of people are going to get accused of AI writing. I talk about concepts like physical objects all the time, and so do people I know IRL, so it might be regional.
Anthropic calls that category mannered prose in their prompting docs. Ask the model to avoid mannered prose and it basically eliminates this type of slop writing.
For me it's the obsession with the universal quantifier. "no model", "every ratio", "every comparison". Probably an effect of training on coding tasks where all cases have to be handled.
This page is so laden with Claude-speak that finding the information in all the noise is hard. The plaintext baseline isn't 100% because a correct answer worded differently is scored as a failure. That's the implementer's phrasing preferences, not proof the phenomenon is emergent.
A useful critique, thanks. The grader has the same threshold for whatever answer it receives and is equally harsh on whichever it grades. Anything above the 1.0 baseline is a parity claim, not better understanding. The decoder expands the cablese without seeing the question, and a separate model instance still answers at parity, so information wasn't lost.
The interesting part for me is BabelTele from June 2026, which showed models encoding text into emoji and symbols at 27.9% of original length with 99.5% fidelity, including cross-model transfer. People tested early GPT-4 the same way back in 2023.
Deeply unserious technology. Can't wait for the article about LLMs scoring 20% better on programming benchmarks if asked to impersonate Kevin from The Office.
Speaking of old timey language, several LLMs were being trained purely on historical text. Have any of them come out yet? The history-llms repo hasn't updated in ten months.
Maybe we should teach them to text like early 2000s teenagers, with SMS billed per 120 characters.
Exact match graders are the real variable here. We've had F1 or a judge model swing passage QA scores by ten points on identical answers.