_has_extension_in and is_sqlite_sidecar each re-implemented the same
rfind('.') / -1 / .lower() suffix extraction, and the write guard used a
third spelling (filepath[filepath.rfind("."):]) for its refusal messages
next to os.path.splitext in the overwrite branch. Add _lower_suffix()
("" when no dot) and route both predicates through it; the write guard
now uses os.path.splitext for every message's displayed extension.
_SQLITE_EXTENSIONS was `{.db,.sqlite,.sqlite3,.db3} & BINARY_EXTENSIONS`,
which silently dropped .db3 (not a binary extension anywhere in the
codebase, so `x.db3-wal` was never a sidecar). Spell the set as the three
members that actually take effect and assert the subset invariant instead
of hiding it behind an intersection. Behaviour is unchanged.
The refactor that moved sidecar detection into binary_extensions dropped the
contributor's unconditional refusal for the db family and made every binary
extension a plain overwrite guard. That is right for the MAIN .db/.sqlite
(text fixtures named *.db exist), but a -wal/-shm/-journal path is never a
legitimate text target: a checkpointed database has no sidecar on disk, so
write_file("kanban.db-wal", text) silently dropped a garbage WAL next to a
live database.
Add is_sqlite_sidecar(path) and refuse such paths in
_check_binary_document_write regardless of is_file(). Scope the marker
stripping in _has_extension_in to SQLite suffixes so report.docx-wal no
longer counts as an opaque document. Hoist the duplicated is_pdf_path call.
Parametrize the write_file WAL test over sidecar existing/absent.
Move the .db-wal/-shm/-journal detection from a write-guard-local regex
into tools/binary_extensions._has_extension_in so has_binary_extension
(read guard AND write guard) agree: a sidecar counts as its database's
extension. read_file now refuses sidecars instead of returning lossy
text, which also gives write_file the no-baseline overwrite refusal.
Behaviour change vs the contributor's commit: creating a NEW .db /
.sqlite file via write_file stays allowed (main allows it and text
fixtures named *.db exist) — only overwriting an existing binary is
refused. The generic has_binary_extension overwrite branch is kept and
merged with the PDF branch because patch has no full-read baseline
check: on main, patch on an existing .db with a matching old_string
rewrites the header in place.
Tests folded into the existing guard test classes (one write_file, one
patch refusal on a real WAL sidecar); the PR's 12-test file and its
private-regex assertions are dropped.
OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file
auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm,
.pps, .pot, .rtf, .epub. Each has the same corruption path: read_file
shows extracted text, model writes it back, container is destroyed.
Flagged by @egilewski on PR #82818 — proven live for .docm (text write
left a non-zip corpse). Added bytes-untouched regression test for .docm.
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx
(and PDF via anydoc) to readable text, so a model plausibly believes it
holds the file's contents and writes the edited text back with
write_file/patch — silently destroying the document container. Proven
live on main: write_file over a valid .docx left a non-zip corpse, and a
text write over an existing .pdf clobbered the %PDF header.
- tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS +
has_opaque_document_extension() + is_pdf_path() (pure string checks)
- tools/file_tools.py: _check_binary_document_write() — opaque container
formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf
rejected only when overwriting an existing regular file (new-PDF
creation stays allowed, matching the upstream split guard). Wired into
write_file_tool and patch_tool (replace + V4A Update/Add headers;
Delete/Move skip the guard since they write no text).
- tests/tools/test_binary_document_write_guard.py: guard unit tests +
end-to-end write_file/patch coverage incl. bytes-untouched assertions.