OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file
auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm,
.pps, .pot, .rtf, .epub. Each has the same corruption path: read_file
shows extracted text, model writes it back, container is destroyed.
Flagged by @egilewski on PR #82818 — proven live for .docm (text write
left a non-zip corpse). Added bytes-untouched regression test for .docm.
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx
(and PDF via anydoc) to readable text, so a model plausibly believes it
holds the file's contents and writes the edited text back with
write_file/patch — silently destroying the document container. Proven
live on main: write_file over a valid .docx left a non-zip corpse, and a
text write over an existing .pdf clobbered the %PDF header.
- tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS +
has_opaque_document_extension() + is_pdf_path() (pure string checks)
- tools/file_tools.py: _check_binary_document_write() — opaque container
formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf
rejected only when overwriting an existing regular file (new-PDF
creation stays allowed, matching the upstream split guard). Wired into
write_file_tool and patch_tool (replace + V4A Update/Add headers;
Delete/Move skip the guard since they write no text).
- tests/tools/test_binary_document_write_guard.py: guard unit tests +
end-to-end write_file/patch coverage incl. bytes-untouched assertions.