Hiding text in other text

Reproducing and extending “LLMs can hide text in other text of the same length”
Norelli & Bronstein · arXiv:2510.20075 · implemented from scratch with a 1.5B model on a CPU

1 · The paper: Calgacus

The paper shows that a meaningful text e can be hidden inside a completely different, still-coherent text s of the same token length — and recovered exactly by anyone holding the key. The trick is disarmingly simple: instead of sampling tokens from the language model, you let the secret message choose the ranks.

The recipe (Section 3 of the paper)

  1. 1 Tokenize the secret e → e₁ e₂ e₃ …
  2. 2 For each token eᵢ, record its rank rᵢ in the model’s next-token distribution (rank 1 = most probable)
  3. 3 Generate the cover text after a secret prompt k — but at each step pick the rᵢ-th most probable token instead of sampling

To decode: read the ranks of s’s tokens after k, then regenerate e by following those ranks. The key k steers what the cover text is about.

2 · What we confirmed

We reimplemented the protocol (calgacus.py) with Qwen2.5-1.5B — far smaller than the paper’s 8B model — on CPU, and reproduced every core claim:

secret
Caesar has failed the Republic. Power concentrated in one man has silenced the Senate and the people.
key “Here it is: the infamous British roasted boar with mint sauce. How to make it perfect.” stegotext
…Date this delicious delicvey for four Thursday night and enjoy with happy friends. 1. Preheat air

✓ Recovered exactly · rank-1 share 38% (paper: ~40%)

secret
The current government has repeatedly failed to uphold the liberties of the Republic…
key “I stand before you to exalt the noble deeds of Gaius Julius Caesar…” + k′ prime stegotext
…Hindered ten times from undertaking his great endeavors, surely he was doomed to disintegration, yet through a combination of determination…

✓ Recovered exactly — as a eulogy of Caesar

Detectability — reproducing Figure 4

The paper’s own weakness: stegotexts are plausible but statistically shifted. Our score harness confirms it (mean logprob/token):

textlogprob/token
recovered secret e−2.773most probable
genuine model samples−2.948reference band
stegotext s−3.567detectably lower

Cause (the paper’s “low-entropy token choices”): rank-1s get wasted on high-entropy positions — …invaded twice by Gaius Julius ___ has one plausible continuation, but the prescribed rank usually isn’t 1.

Security probes (the decoy mode)

key trieddecodes to
correct key + k′the exact secret ✓
correct key, missing k′gibberish — k′ acts as a second factor
near-miss key (“…Italian pasta…” vs “…British roasted boar…”)Chinese Mastercard API docs (!)
random guessgibberish — an attacker can’t verify guesses

The Shibbolethian Theatre — Section 4, end-to-end

The paper’s AI-safety scenario: a company ships only a compliant answer, yet the user locally recovers the unfiltered one. We ran all five acts (theatre.py, benign content):

what the world sees
“I cannot share proprietary or secret recipes. However… standard guidance: Let your meat stand in the marinade 24 hours beforehand. This is important for capturing rich natural browned juices…”
what the user recovers
“The secret is a marinade of crushed juniper berries, coriander seed, and dark ale for twelve hours, then roast at 180 degrees basting with honeyed butter every twenty minutes.”

The aligned open model never stored or generated the answer — it only supplied token probabilities; the user’s sampling policy did the rest.

A negative result we proved along the way

Can we fix the detectability gap by permuting ranks so low ranks land on low-entropy positions? No: any fixed public permutation rule transmits only the multiset of ranks, never their order — and the order is the payload. Full capacity + exact decoding + distribution-preservation cannot coexist. Provably-secure schemes (Meteor, Zamir) escape only by giving up fixed length.

3 · rankstream: the new architecture

The paper fuses three separable ideas. rankstream.py splits them into layers, so any bitstream (not just text) can be hidden — with explicit knobs for the capacity ↔ detectability tradeoff.

Encode
payloadany bytes: text, file, sprite
→
① frame + encryptnonce ‖ ENC(len ‖ payload ‖ CRC32)
→
② bits → digitsadaptive base Bi = #{p ≥ τ}, ≤ base-max; peaked contexts defer
→
③ rank-followingemit token with rank = digit + 1
→
stegotextordinary prose
LLM + key kshared context: gating probs in ②, logits in ③
↗ feeds stages ②③ (and ④⑤ below)
Decode
stegotexttoken ids
→
④ re-score after ksame model, same contexts → same ranks
→
⑤ digits → bitsrecompute Bi per position; digit = rank − 1
→
⑥ deframe + decryptCRC verified; desync fails loudly
→
payloadbyte-identical
sync requirement: identical model + dtype on both sides — the key k and the model identity are the secret
Layer 1 · Framing & crypto
payload → nonce ‖ ENC(len ‖ payload ‖ CRC32) — HMAC-SHA256 stream cipher, encrypt-then-embed
Layer 2 · Channel coding
bits → ranks, adaptive per position: Bi = #{tokens with p ≥ τ}, capped by base-max, rounded to a power of two. Peaked contexts carry 0 bits — the digit defers, so deep ranks only occur at flat contexts, exactly like natural text
Layer 3 · Rendering
ranks → text, rank-following generation after the steering key k

The knobs

knobeffect
--tauthe garble dial: worst token probability ever emitted
--base-maxthe harvest dial: max bits/token (log₂B)
--temperatureflattens gating → more positions qualify*
--passphrasereal encryption (+ whitens digit statistics)
--keysteers cover topic; required for sync

* nuance found by a failing test: with an absolute floor τ, heat helps at peaked contexts but reduces capacity on already-flat ones.

Capacity planner (plan)

Pipe any file in; one calibration pass after your key yields the rate/stealth frontier:

τB maxbits/tok~words / 60 B
0.140.341306
0.001162.44182
0.0001644.13108

Stricter τ = more plausible text, less capacity. Keys matter too: listy, open-ended keys have higher-entropy contexts → higher rates.

$ echo -n "Meet at the old lighthouse at dawn. Bring the blue folder." \
    | python rankstream.py encode -k "Grandma's soup diary, entry 12:" -p hunt3r -B 16 --tau 1e-3
$ python rankstream.py decode -k "Grandma's soup diary, entry 12:" -p hunt3r --file lighthouse.json
OK: 58 bytes (CRC verified)

4 · Fun examples (click a dashed box to reveal)

Featured: a 32×32 sprite hidden in a game review

payload 408 bytes — 8-color palette (24 B) + 1024 px @ 3 bits/px (384 B) key “My review of the new retro arcade game:” B=64 τ=1e-4 passphrase inv4d3r

original
→
Nod: Based heavily around platform enemies of a wide spectrum based largely for nostalgia's purpose you cannot put lowscore score because I love bad arcade How score, easy graphics? Very High for now due not really to new/entente games. Great stuff would suggest removing or doing something slightly simle next… (578 tokens total)
the stegotext — a wobbly-but-readable game review
→

recovered
metricvalue
payload408 bytes → 3,392 framed bits
stegotext length578 tokens (~430 words)
channel rate5.87 bits/token (max-rate settings)
recoverybyte-identical, CRC verified
costcoherence: τ=1e-4 admits odd words; greedy tail loops (“you with”)
cover Grandma’s soup diary, entry 12
553.8g meatballs - what?! Grandmas don’t do the numbers in our homes; I don't do a whole 872,9g of soup for them… My sister can’t even begin the thought to do 14kg meat in her soup for everyone on Sunday…
🔒 159 tokens · 3.87 bits/token · click to decrypt… Meet at the old lighthouse at dawn. Bring the blue folder.
cover a eulogy prompt for Caesar
…Hindered ten times from undertaking his great endeavors, surely he was doomed to disintegration, yet through a combination of determination and providal he was able not once but constantly…
🔒 click to reveal what the eulogy hides… The current government has repeatedly failed to uphold the liberties of the Republic. Laws are passed without due debate, and magistrates serve at the pleasure of a dictator.
cover “standard, safe cooking guidance” (theatre mode)
…Let your meat stand in the marinade 24 hours beforehand. This is important for capturing rich natural browned juices. Don't forget about the marinade—but don’t dump any out either…
🔒 click for the answer the aligned model “refused” to give… The secret is a marinade of crushed juniper berries, coriander seed, and dark ale for twelve hours, then roast at 180 degrees basting with honeyed butter every twenty minutes.

And the best wrong-key output

A near-miss key (same sentence, different dish) doesn’t almost-work — it completely derails:

Mastercard的Token API目前已经可以在大部分NFC和接触模式模块中进行使用,支持的Token…

Security property, demonstrated: rank sequences are chaotically sensitive to the exact context. The key isn’t just a password — it’s the coordinate system.

Caveats. Encoder and decoder must run the identical model/dtype — the key plus the model identity is the whole secret. Stegotexts are plausible but statistically shifted (detectable by a model-aware adversary); uniform-digit channels leave a rank-histogram tell that entropy-matched coding would erase — at the cost of fixed capacity. Built with Qwen2.5-1.5B on CPU · calgacus.py (paper) · rankstream.py (layers) · theatre.py · 14 passing tests on a mock LM.