OtherShrezzing 7 hours ago

This page is (somewhat ironically) so extremely laden with Claude-speak that it's difficult to find the information in all the noise. But once you've waded through everything, you see these facts:

>What the test measures: A model is given a passage and a fixed set of questions with short, checkable answers — a date, a name, a count.

So, a model is given content which is especially amenable to compression, and asked to reproduce it under certain constraints, like...

>Why isn’t the plaintext baseline 100%? Answering questions about an uncompressed passage in plaintext scores ~91%.... a correct answer worded differently scores as a [failure]

Models can (and do) give objectively correct answers, but are penalised for not having some kind of omniscient knowledge of the implementer's phrasing preferences.

If this phenomenon is emergent in models, this benchmark is not proof of it in any meaningful way.

  • andai 6 hours ago

    For me the really interesting part was:

    > Haven’t we seen LLMs do this already?

    > Yes, BabelTele (arXiv, June 2026) demonstrated that LLMs can encode text in compact, non-standard forms — omnilingual word fragments, symbols, emoji — that other models recover with high fidelity (99.5% semantic fidelity at 27.9% of original length, by their metrics), including cross-model transfer, agent memory, and multi-agent communication. It proves the general phenomenon: human readability is not a requirement for model-to-model text.

    I remember people testing early GPT-4 (2023?) in similar ways, to compress text, it would emit a string of strange text, Unicode, emojis, but was able to decode the compressed version very reliably.

    This seems to cut usage by another ~50%, at the cost of being incomprehensible to humans.

    • apefulsin 6 hours ago

      Once, a GAN model that was trained to convert between satellite images and drawn maps was caught encoding the original satellite image in imperceptible dots

      The field of ML is Goodhart's law reified. We might have temporarily forgotten some of the basics of the field amidst this LLM craze.

    • Theory42 5 hours ago

      Yeah! LLM compression is a spectrum between 'normal stuff we can read' and vectors. Depending on trust in the models and the desire for compression, there's a choice to be made on which formats you want to allow. Pretty interesting stuff.

    • HPsquared 4 hours ago

      Good for chain of thought, perhaps?

    • andai 6 hours ago

      [dead]

  • Theory42 6 hours ago

    A useful critique, thanks.

    I would say that the grader has the same threshold for whatever answer it receives, and is equally harsh on whichever it grades. Any scores above the baseline (1.0) are really claims about parity, rather than better understanding in the compressed format.

    The decoder step is a model expanding the cablese to regular text, not having seen the initial question. A separate model instance then reads that regular text and answers. And the result is still at parity with the plaintext record.

    Had cablese knocked out information, that wouldn't have been the result, would it?

  • mattmanser 2 hours ago

    Ha, even you quoting the Claude-ese made me zone out of your comment and switch tabs off hacker news. As soon as I saw "What the test measures".

    I only realized why I'd switched tab after I'd done it!

Solomet 7 hours ago

Newest LLM writing tell: Concepts are described in terms normally more appropriate for physical object.

> A lab that suppresses it in a frontier model just moves the advantage to open models that still _carry_ it

> they carry no signal about which is better

> where your workload _sits_ on that frontier should pick the point

> and no model _sits_ in the judge’s seat

> every ratio _sits_ at 0.99–1.10

Many many more examples of "sit"

> Every comparison in this post "holds" the questions

I have been seeing this a lot in my recent work with LLMs and it is quite frustrating. Even more frustrating is how frequently it uses low-signal terms for things unnecessarily. These 'physical object' terms are one example but at times it really seems that they 'preserve effort' by choosing a less descriptive term because it 'fits'

I have also caught it replacing descriptive terms with more vague ones for no discernible reason other than laziness.

"Minimize ambiguity" has been my go-to instruction as of late when the agent drifts back towards vague terms and lack of specificity.

  • ipdashc 5 hours ago

    For me it's the obsession with the universal quantifier. Even in these examples: "no model", "every ratio", "every comparison". They love emphasizing that everything in a set meets some condition. I assume it's an effect of being trained on coding tasks where they need to make sure that all cases are handled.

  • Hugsbox 7 hours ago

    Oh crap, if these are the new LLM tells then a lot of people are going to start accusing me of AI writing...

    I have a strong tendency of talking about concepts like they're physical objects. A lot of the people I know IRL do too, so it might be a regional thing idk.

    • lopis 7 hours ago

      I think these are growth pains. As LLM start to grasp new figures of speech, it sounds weird at overuse at first, until it finds a balance.

  • ruuda 5 hours ago

    By now I am allergic to the word "carry", I just cannot continue reading any more.

  • elzbardico 4 hours ago

    Business folks speach have this annoying tendency too.

    I have the distinct impression that MBA and salesman folks think that adequate mathematical terminology is somewhat less "macho", and this impression is re-inforced by the fact that they also love military-adjacent terms and analogies.

  • criley2 7 hours ago

    >"Minimize ambiguity" has been my go-to instruction as of late when the agent drifts back towards vague terms and lack of specificity.

    Anthropic has called the greater category containing this type of writing "mannered prose" https://platform.claude.com/docs/en/build-with-claude/prompt...

    If you ask the models to avoid mannered prose (or use their extended prompt), it basically eliminates all of this type of slop writing.

    Here's a de-slopped example.

    > Write Like It's 1866: LLMs Relearn Telegraphese

    > Adding one sentence to a prompt, telling the model to write like a telegram, cut its output tokens by 40–49%. The sentence asks it to drop articles and filler but keep every fact. Models from four different labs then answered questions from that compressed text as accurately as from normal English. So when one model writes something for another model to read, you pay about half as much for the output. This post introduces the Telegraph Test, a benchmark that measures how well a given model does this.

  • Theory42 7 hours ago

    Cool story bro. Maybe you could engage with the content? I'm an actual person.

    • cyclopeanutopia an hour ago

      If you call your ideas "content", then there is no point in reading it anyway.

    • andai 6 hours ago

      Wait, you wrote this manually?

      • Theory42 6 hours ago

        I wrote it with the help of an LLM, then edited it, then wrote some more, then edited that. Took about a week to format. It contains my thoughts. Brave new world, I know. I'm finding it ironic that the same crowd that is merging AI code all day is so allergic to AI assistance in prose...do all your work using this new godlike technology, but when it comes time to share, whip out your fountain pen or else.

        • boonzeet 6 hours ago

          The problem is that the prose becomes difficult to read because of its lexical quirks and verbiage. One AI generated piece of prose in isolation, fine; but people have developed a 'smell' of it and a mental association with low-quality work.

        • OrderlyTiamat 2 hours ago

          I think AI assistance in writing is fine for what it's worth. Perhaps you'd like https://sockpuppet.org/blog/2026/09/17/how-to-write-with-an-... or https://www.seangoedecke.com/how-i-use-llms/#proofreading-fo...?

          However, I'm as allergic to slop in my code as in my reading. I'd hit "request changes" on this blog post- LLM assistance is as irrelevant to that as it is in a PR.

          I'll commit to the bit: here's a PR review on your blog post. Just my humble opinion and I'm no writer myself, so feel free to disregard all of it. It was written by hand, minor rewrite according to AI critique :)

          ---

          > Instructed to answer in cablese — the telegraph operators' compressed dialect, models elicit 40–49% fewer billed output tokens on the API's meter, and models across four families still recover the information at full fidelity.

          suggestion: that's a run on sentence, with a "—" signifying a new sentence, but this second sentence doesn't have a clear flow. Try vocalising this sentence, where are your breathing pauses? I can't vocalise it clearly. It's a very minor thing! Try vocalising this minor rewrite: "Instructed to answer in cablese — the telegraph operators' compressed dialect — models elicit 40-49% fewer billed output tokens. Models across four families still recover the information at full fidelity."

          That's much easier to vocalise. Usually, that makes it easier to read, too.

          > The Telegraph Test benchmark measures how well models compress using this technique as well as how much information they can retreive from it afterward.

          nit: reword. "how well models compress" is a bit awkward here. something like "The Telegraph Benchmark measures the model's ability to compress, as well as...". To me "Telegraph Test benchmark" is a confusing term, "Telegraph Benchmark" is clearer, and still terse(er!)

          praise: otherwise this is a good abstract-like introduction.

          nit: You did clearly edit this by hand, because you misspelled "retreive".

          > Cross-family matrix (readers = foreign models answering from GLM-5.3-Flash’s records; writers = GLM-5.3-Flash answering from theirs) :

          suggestion: add a paragraph. This is not a good introductionary text, what kind of questions did you test on, what's the goal here?

          > GLM-5.3-Flash itself: 48.4% savings with the lowercase instruction, in-family recovery 1.09. No comparison in the matrix favors plaintext; every ratio sits at 0.99–1.10.

          question: why start with "<one of the models> itself?" It's unclear why you're talking about that one in particular here, and makes it hard to follow your point.

          > The condition ladder — same questions, one variable at a time:

          question: what is a condition ladder?

          > The register is not a construct we invented (LLMs were handed compressed records cold and read them at parity); the capability was already in the weights, inherited from a century and a half of people writing under metered bandwidth.

          suggestion: rephrase. What's "the register", as in the tone the LLMs speak in? explain your terms.

          > Every model tested can do this.

          suggestion: rephrase, that's not a grammatically correct sentence. e.g. "all models we tested can do this", or "every tested model can do this" if you wish terseness.

          One such sentence is obviously not a problem at all! but too many, and you'll lose readers.

          > Notice what this adds up to: Result:

          praise: good centerpiece, that's your central thesis

          > No new hardware, no training, no API change, one sentence of instruction.

          nit: very AI coded language, "no <x>, no <y>, ..." is cliché. Not a problem obviously, but I thought I'd bring your attention to it.

          ---

          That's sorta where I lost interest in this PR review bit. the point here is that it's harder to read and engage your content. Also, if it looks too AI, you'll lose readers who assume you didn't put effort in.

    • soleman 7 hours ago

      Why not write your article in the same telegraphese you preach? Hilariously ironic to use verbose AI writing for this.

      Also the site background is AI slop which makes for terrible contrast with the text.

      • Theory42 7 hours ago

        I use the tools I study, nothing more or less.

        • jchw 7 hours ago

          Since when is writing blog posts themselves part of "studying"?

          • SoftTalker 3 hours ago

            Blogging about something you're studying? Seems like it could be effective. It presents your understanding of the subject to the world, and invites critique.

            • jchw 2 hours ago

              I'm starting to think the abstraction involved in internet communication is really confusing some people. When you copy and paste LLM output verbatim into a text box and send it to someone, that's weird, right? You're sending that as if you just said it. But you didn't. That's what this here is. Even just adding a "this was AI generated" disclaimer doesn't make it less strange.

              If this was the real world, the physical equivalents of this would be utterly psychotic and nobody would hang out with you. "Hey, are you just typing everything I say into ChatGPT and reading out what it says?" "No, not completely. I'm changing a few of the words. Plus it's the future, everyone is using LLMs!"

              If you're someone who insists on publishing almost entirely AI-generated blog posts, I think what you need to do is give your AI agents wordpress instances because it seems like they're the ones doing most of the blogging and who people who find the posts interesting might want to follow.

    • Solomet 6 hours ago

      > Instructed to answer in cablese — the telegraph operators' compressed dialect (drop the articles and filler, keep every fact) — a single one-sentence instruction, no examples and no codebook, elicits 40–49% fewer billed output tokens on the API's own meter, and models across four families still recover the information at full fidelity. For machine-to-machine traffic, that is half the output bill at any major API, today. The Victorian economics of the cable, reborn as token economics: the Telegraph Test benchmark.

      I tried but this first paragraph seems like it was almost purposefully obfuscated. Typical LLM-written content that meanders around a bit and stops when it seems to have emitted enough words. The reader is left to assemble meaning from the trace of thought it did not go back over to revise.

      I think you have a good point to make but the writing is really difficult to get past.

netsharc 6 hours ago

The Cablese/Telegraphese is more interesting than the use in LLM. DuckDuckGo'ed "paromella":

https://en.wikipedia.org/wiki/Commercial_code_(communication...

Some codes I found interesting:

> INSANE - at what price, free on board and freight, can you offer us cotton for shipment by steamer sailing this week?

> COGNOSCO - dining out this evening, send my dress clothes here

Useful codeword!

> ANNOSUS — Confined yesterday, Twins, both dead, Mother not expected to live

How often did that one come into use??

z2 8 hours ago

From recent ChatGPT (GPT5.6) conversations where I've seen occasional reasoning leaks into the UI, it's clear that something like this is already implemented, and I'd speculate that this is the majority of recent claims of less token usage. Not sure if they are literally prompting for cablese of course.

"Need check output vs prev. Ran script, results fine, need prep next step. Ready? Go."

  • lxgr 3 hours ago

    This is their CoT reasoning trace. (Why you see it: Models are supposed to delimit "actual" CoT, their user-visible summary/"cleaned CoT", tool calls, and user output via delimiters, but don't always do it perfectly.)

    Rewarding terse CoTs during training seems like a no-brainer, as long as it doesn't impact capabilities, so I suspect the style is completely emergent and probably what you get when you implement a dual goal of terseness and capabilities while still punishing completely non-human-readable CoT. (Failing to do the last part would probably have models speak in ominous Unicode glyphs in no time.)

  • Theory42 7 hours ago

    I suspect you might be right. My conclusions from the digging are that this kind of compression works best with settled instructions/data for machine to machine talk. For something like OpenClaw (which I use a lot), that might mean the AGENTS.md, TOOLS.md, etc. Compression there would free up the context for the agent.

  • andai 6 hours ago

    Would be pretty funny if OpenAI's models are so token efficient now due to hidden caveman prompt.

    • tk75x an hour ago

      Reminds me of how we used to search with keywords that ended up sounding like caveman speak, then there was "natural language search", and now the AI is talking to itself in caveman speak again.

  • _fw 7 hours ago

    I can confirm I’ve seen this with Deepseek V4.1, I imagine it’s with other models too.

1vuio0pswjnm7 34 minutes ago

Maybe agents can communicate via Morse code through port knocking, stick tables or something similar

alexpotato 7 hours ago

Actor to Winston Churchill:

"Show premiere Oct 10th STOP Bring a friend STOP If you have one STOP"

Winston Churchill to actor:

"Can't make premiere STOP Will come to second showing STOP If there is one STOP"

bilater 2 hours ago

Actually, I think the lesson from this article is that a lot of these little hacks we’re doing around context compaction, AGENTS.md, system guides, and other harnessy ways of saving tokens are going to go away fairly quickly, just as all the tricks of the telegraph era went away once communication got cheap enough that it was easier to just speak in plain language.

novideonoradio 7 hours ago

Deeply unserious technology. Can't wait until an article about LLMs performing 20% better on programming benchmarks if asked to impersonate Kevin from The Office.

erelong 6 hours ago

Yeah I've thought of speaking like a caveman before to AIs but also maybe we could communicate more simply with people; ironically the article could be rewritten in telegraphese or caveman-speak

Would be nice to see language engineered to communicate more simply (like the idea of -- not necessarily implementation -- simple Wikipedia)

Also articles like this sprawl a bit and idk how to even make them easier to read (maybe AI has ideas to make reading and writing simpler)

swiftcoder 6 hours ago

Maybe we should teach them to text like early 2000's teenagers, with SMS billed by the 120 chars...

  • lxgr 3 hours ago

    160 characters in latin-ish alphabets, or 70 in Unicode. (Hard to forget when you had to pay for every single text from your pocket money :)

hnd9q09qk4 6 hours ago

Exact match graders are the real variable here, we had F1 or a judge model swing passage QA scores by ten points on identical answers.

  • Theory42 6 hours ago

    A good point. I initially had a model for a judge, but it seemed to give very lenient scores. I'm open to learning about how best to benchmark the phenomenon, though.

klaff 6 hours ago

Why the terrible AI image up top? Non-functional telegraph key, telegram that looks nothing like a real one, infant-sized bowler. I guess we're past rampant nonsense words, so progress?

jubilanti 8 hours ago

Just another AI slop version of the old 'caveman' dialect.

  • Theory42 8 hours ago

    Not quite, and its addressed in the post. Caveman was cool, but hand wavy. I've measured the compression of Cablese, determined where it does the most good, and also that it is already baked in to most model family's training, which saves instruction tokens.

    • 6031769 6 hours ago

      > Caveman was cool, but hand wavy.

      Cavey-wavy. Ahem.