Docs / Commands

Read documents and media

Updated

TL;DR — `tokenade read <file>` extracts text from PDF, Office, OpenDocument, EPUB and more; images go through OCR, audio and video through subtitles or a local transcript. Add `--prompt "q1, q2"` to get only the passages that answer.

tokenade read is the single entry point for content: give it a file (or - for stdin) and it detects the format and returns text your agent can use. For documents and media that means extracted text, OCR or a transcript instead of raw bytes. Add --prompt and you get only the passages that answer your questions.

Usage

tokenade read <file|-> [--prompt "q1, q2"] [--lines A-B] [--cmd <command>] [--frames [N]]
FlagEffect
--prompt "q1, q2"Return only the passages that answer. Several comma-separated questions in one call. Also spelled --question, -q
--lines A-BRead only that region, 1-based and inclusive. A- runs to the end, -B from the start, A is one line. A range past the end is refused
--cmd <command>Name the command that produced the text, so the right per-command compactor is used
--frames [N]Video only: add deduplicated still frames (default 8, at most 32)

Stdin needs -: curl -s https://api.example.com/x | tokenade read -.

Supported documents

FamilyExtensions
PDF.pdf (scanned pages are OCR'd when possible)
Word.docx .docm .dotx .dotm .doc .dot .rtf
Excel.xlsx .xlsm .xltx .xltm .xls .xlsb
PowerPoint.pptx .pptm .potx .potm .ppsx .ppsm .ppt .pps .pot
OpenDocument.odt .ott .ods .ots .odp .otp .odg .otg, flat .fodt .fods .fodp .fodg
E-books and mail.epub .fb2 .msg
Images.png .jpg .jpeg .gif .webp .bmp .tif .tiff .ico and other common formats
Audio.mp3 .wav .m4a .flac .ogg .opus .aac and other common containers
Video.mp4 .mkv .mov .webm .avi and other common containers

Structured text (JSON, YAML, CSV, logs, diffs, stack traces, notebooks…) goes through the same command and is compacted by format. Source code is passed through unchanged.

Ask, don't read

A short document is returned whole. A long one is better asked:

tokenade read contract.pdf --prompt "termination notice period, governing law, liability cap"

You get each answer under its own heading, in one round-trip. Several questions in one call is the cheapest option in tokens.

Each comma-separated part should have at least 3 words. A shorter part is merged into its neighbour and the output tells you so:

[tokenade] "due date" is too short to stand on its own as a question (1 of the 2 comma-separated parts fall under 3 words), so it was folded into its neighbour …

For a big text file you need piece by piece, use --lines:

$ tokenade read --lines 1-3 PATCHNOTES.md
[tokenade: lines 1-3 of 2108]
# Tokenade — what's new
## 1.2.2

Images

With tesseract installed, read returns the text the image shows:

$ tokenade read dashboard.png
[tokenade:image 890x817] text read by OCR (tesseract) — recognition can misread characters; the image itself is at dashboard.png:
…

Set the OCR language with TOKENADE_OCR_LANG (default eng; any installed tesseract language pack). Scanned PDF pages are OCR'd too when pdftoppm and tesseract are both installed (first 20 marked pages).

When your agent reads an image itself on Claude Code, the Read hook serves a downscaled copy if the long edge is over 1024 px (the default). Set TOKENADE_IMAGE_MAX_EDGE=0 to keep full resolution.

Audio and video

A media file is never decoded as text. read returns what the file is (duration, codecs, tracks, from ffprobe) and what is said in it, from the cheapest source available:

  1. a subtitle file next to the media (talk.mp4 → talk.srt, .vtt, .ass, .ssa);
  2. a subtitle track embedded in the container (needs ffmpeg);
  3. a local transcript with whisper, or whisper-cli plus ffmpeg.

Transcription runs only from the tokenade read command, never inside an agent hook: a hook only returns the short description. Transcripts are cached by content, so a renamed or copied file is not transcribed again.

VariableDefaultEffect
TOKENADE_WHISPER_MODELbaseWhisper model: tiny for speed, small or medium for accuracy
TOKENADE_WHISPER_TIMEOUT_SECS900Wall-clock limit for one transcription

--frames extracts evenly spaced still frames (deduplicated) into the stash so the agent can look at them. Each frame costs vision tokens when read, which is why the default stays at 8.

Optional tools

None of these are required. Without them read still identifies the file and refuses to inflate it, and names the binary that would unlock more.

ToolUnlocks
tesseractOCR of images and scanned PDF pages
pdftoppmRendering scanned PDF pages for OCR
ffprobeDuration, codecs and tracks of media files
ffmpegEmbedded subtitles, frames, whisper-cli input
whisper or whisper-cliLocal transcripts

Exit codes and gotchas

  • Exit 0 on success; 2 for a usage error such as a --lines range past the end of the file.
  • Secrets are redacted in every output (for example Routing: <redacted>), whichever compactor was picked.
  • If compaction would make the output larger than the input, the raw bytes come back with a passthrough_inflated label.
  • To force a specific command's compactor on captured output, use --cmd, or run the producing command through tokenade wrap.