Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
The test trim dropped the headline case: text written over an EXISTING .db. Sidecar paths return at is_sqlite_sidecar before the overwrite branch, so a mutant that drops has_binary_extension from the overwrite condition survived every remaining test. Add a third parametrize case that targets the held-open WAL database file itself and asserts the binary refusal with bytes unchanged.
Same commit: the _check_binary_document_write docstring now names SQLite sidecars (always) and every BINARY_EXTENSIONS suffix (on overwrite), and the suffix is computed once at the top instead of three times.
_make_wal_db claimed sqlite keeps the -wal until checkpoint, but SQLite
deletes -wal/-shm on last-connection close, so the fake-bytes fallback was
what actually ran. Make the fixture a context manager that holds the
connection open across the tool call so the sidecar is a real WAL, and drop
the PRAGMA integrity_check assertion (SQLite ignores a garbage WAL, so it
passed either way); keep the byte-identity check.
In the patch test assert "binary" in the error so the no-baseline guard
cannot mask a regression in sidecar detection.
The refactor that moved sidecar detection into binary_extensions dropped the
contributor's unconditional refusal for the db family and made every binary
extension a plain overwrite guard. That is right for the MAIN .db/.sqlite
(text fixtures named *.db exist), but a -wal/-shm/-journal path is never a
legitimate text target: a checkpointed database has no sidecar on disk, so
write_file("kanban.db-wal", text) silently dropped a garbage WAL next to a
live database.
Add is_sqlite_sidecar(path) and refuse such paths in
_check_binary_document_write regardless of is_file(). Scope the marker
stripping in _has_extension_in to SQLite suffixes so report.docx-wal no
longer counts as an opaque document. Hoist the duplicated is_pdf_path call.
Parametrize the write_file WAL test over sidecar existing/absent.
Any error would also be produced by the no-baseline overwrite guard on
main, so the test could not detect a regression of the sidecar
detection. Asserting the binary-file refusal makes it fail when
binary_extensions stops recognising .db-wal.
Move the .db-wal/-shm/-journal detection from a write-guard-local regex
into tools/binary_extensions._has_extension_in so has_binary_extension
(read guard AND write guard) agree: a sidecar counts as its database's
extension. read_file now refuses sidecars instead of returning lossy
text, which also gives write_file the no-baseline overwrite refusal.
Behaviour change vs the contributor's commit: creating a NEW .db /
.sqlite file via write_file stays allowed (main allows it and text
fixtures named *.db exist) — only overwriting an existing binary is
refused. The generic has_binary_extension overwrite branch is kept and
merged with the PDF branch because patch has no full-read baseline
check: on main, patch on an existing .db with a matching old_string
rewrites the header in place.
Tests folded into the existing guard test classes (one write_file, one
patch refusal on a real WAL sidecar); the PR's 12-test file and its
private-regex assertions are dropped.
OPAQUE_DOCUMENT_EXTENSIONS was missing 10 extensions that read_file
auto-extracts via anydoc: .docm, .xlsm, .xlsb, .pptm, .ppsx, .ppsm,
.pps, .pot, .rtf, .epub. Each has the same corruption path: read_file
shows extracted text, model writes it back, container is destroyed.
Flagged by @egilewski on PR #82818 — proven live for .docm (text write
left a non-zip corpse). Added bytes-untouched regression test for .docm.
Port from nearai/ironclaw#7109: read_file auto-extracts .docx/.xlsx/.pptx
(and PDF via anydoc) to readable text, so a model plausibly believes it
holds the file's contents and writes the edited text back with
write_file/patch — silently destroying the document container. Proven
live on main: write_file over a valid .docx left a non-zip corpse, and a
text write over an existing .pdf clobbered the %PDF header.
- tools/binary_extensions.py: OPAQUE_DOCUMENT_EXTENSIONS +
has_opaque_document_extension() + is_pdf_path() (pure string checks)
- tools/file_tools.py: _check_binary_document_write() — opaque container
formats (doc/docx/xls/xlsx/ppt/pptx/odt/ods/odp) always rejected; .pdf
rejected only when overwriting an existing regular file (new-PDF
creation stays allowed, matching the upstream split guard). Wired into
write_file_tool and patch_tool (replace + V4A Update/Add headers;
Delete/Move skip the guard since they write no text).
- tests/tools/test_binary_document_write_guard.py: guard unit tests +
end-to-end write_file/patch coverage incl. bytes-untouched assertions.