Commit Graph

2 Commits

Author SHA1 Message Date
kshitij
5c988e2461 fix: cover remaining anydoc-extracted container formats
OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file
auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm,
.pps, .pot, .rtf, .epub. Each has the same corruption path: read_file
shows extracted text, model writes it back, container is destroyed.

Flagged by @egilewski on PR #82818 — proven live for .docm (text write
left a non-zip corpse). Added bytes-untouched regression test for .docm.
2026-08-15 02:51:59 +05:30
Teknium
6d51c831eb fix(file_tools): refuse plain-text writes that corrupt binary documents
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx
(and PDF via anydoc) to readable text, so a model plausibly believes it
holds the file's contents and writes the edited text back with
write_file/patch — silently destroying the document container. Proven
live on main: write_file over a valid .docx left a non-zip corpse, and a
text write over an existing .pdf clobbered the %PDF header.

- tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS +
  has_opaque_document_extension() + is_pdf_path() (pure string checks)
- tools/file_tools.py: _check_binary_document_write() — opaque container
  formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf
  rejected only when overwriting an existing regular file (new-PDF
  creation stays allowed, matching the upstream split guard). Wired into
  write_file_tool and patch_tool (replace + V4A Update/Add headers;
  Delete/Move skip the guard since they write no text).
- tests/tools/test_binary_document_write_guard.py: guard unit tests +
  end-to-end write_file/patch coverage incl. bytes-untouched assertions.
2026-08-15 02:51:59 +05:30