feat(desktop): group a comment batch by page region so it lands as a few tasks

Twenty-three comments arrived as twenty-three flat blocks, so the agent made
twenty-three todos and ground through them one at a time. They now arrive
grouped by where they sit in the page, with a line telling the agent to work
the groups rather than the comments.

The renderer groups on structure, not meaning. Whether a comment is a UI nit
or a functional bug is a judgment only the model can make, and prose-matching
it here would be wrong constantly; which pins share a DOM subtree is something
the selector already answers. That split is also the one that makes parallel
work safe — grouping by theme instead ("all the spacing ones") cuts across the
same components and puts several workers in the same files, so the guidance
says to hand out whole groups and never to regroup by theme.

Grouping compares ancestor paths, so a heading and a paragraph in one card
stay together instead of becoming two singletons. Depth is derived rather than
tuned: descend the shared prefix until it stops being shared, then sub-split
any group still holding more than a third of the batch — without that pass a
normal page buries every section under `main`. Batches under four comments,
and batches that all land in one region, stay flat.

Grouping is advice in the prompt, never an action: the renderer does not spawn
or delegate anything. That stays the agent's call.
This commit is contained in:
Brooklyn Nicholson
2026-09-02 12:12:32 -05:00
committed by brooklyn!
parent e4bda3ff77
commit 06a4f4ab31
6 changed files with 476 additions and 6 deletions

View File

@@ -0,0 +1,165 @@
import { describe, expect, it } from 'vitest'
import { annotateSplitDepth, groupAnnotations } from './group'
import type { ComposerReadyAnnotation } from './pack'
function item(number: number, selector?: string): ComposerReadyAnnotation {
return {
imageDataUrl: '',
note: `note ${number}`,
number,
prompt: `Comment ${number}`,
identity: selector ? { css: {}, html: '', selector, tag: 'div', text: '' } : undefined
}
}
describe('annotateSplitDepth', () => {
it('splits at the shallowest region where the comments disagree', () => {
expect(annotateSplitDepth(['body>main>div.header>h1', 'body>main>div.header>p', 'body>main>div.footer>a'])).toBe(3)
})
it('does not split siblings inside one component', () => {
// Differ only at the leaf, so the ancestor paths are identical: the depth
// runs to the end of the shared path and both land in the same group.
const selectors = ['body>div.card>h1', 'body>div.card>p']
expect(annotateSplitDepth(selectors)).toBe(2)
expect(groupAnnotations(selectors.map((selector, index) => item(index + 1, selector)))).toHaveLength(1)
})
it('does not split when every comment is on the same element', () => {
expect(annotateSplitDepth(['body>div.card', 'body>div.card'])).toBe(1)
})
it('separates a container comment from comments nested inside it', () => {
const depth = annotateSplitDepth(['body>main', 'body>main>div.a>span', 'body>main>div.b>span'])
expect(depth).toBe(2)
})
it('handles a single selector and an empty batch', () => {
expect(annotateSplitDepth(['body>div.only'])).toBe(1)
expect(annotateSplitDepth([])).toBe(0)
})
})
describe('groupAnnotations', () => {
it('gathers comments on the same region and separates different regions', () => {
const groups = groupAnnotations([
item(1, 'body>main>section.hero>h1'),
item(2, 'body>main>section.pricing>button'),
item(3, 'body>main>section.hero>p'),
item(4, 'body>main>section.pricing>span')
])
expect(groups).toHaveLength(2)
expect(groups[0]?.label).toBe('section.hero')
expect(groups[0]?.items.map(entry => entry.number)).toEqual([1, 3])
expect(groups[1]?.label).toBe('section.pricing')
expect(groups[1]?.items.map(entry => entry.number)).toEqual([2, 4])
})
it('produces groups whose subtrees do not overlap, so they can run in parallel', () => {
const groups = groupAnnotations([
item(1, 'body>main>section.hero>h1'),
item(2, 'body>main>section.pricing>button'),
item(3, 'body>main>section.faq>li')
])
const keys = groups.map(group => group.key)
const overlapping = keys.filter(key => keys.some(other => other !== key && other.startsWith(`${key}>`)))
expect(keys).toHaveLength(3)
expect(overlapping).toEqual([])
})
it('keeps area pins in their own trailing group rather than guessing a subtree', () => {
const groups = groupAnnotations([item(1, 'body>main>div.a>h1'), item(2), item(3, 'body>main>div.b>h1'), item(4)])
const loose = groups[groups.length - 1]
expect(loose?.key).toBe('')
expect(loose?.label).toBe('')
expect(loose?.items.map(entry => entry.number)).toEqual([2, 4])
})
it('returns one group when every comment lands in the same region', () => {
const groups = groupAnnotations([item(1, 'body>div.card>h1'), item(2, 'body>div.card>p')])
expect(groups).toHaveLength(1)
})
it('orders groups by first appearance so the numbering still reads in click order', () => {
const groups = groupAnnotations([
item(1, 'body>main>div.b>h1'),
item(2, 'body>main>div.a>h1'),
item(3, 'body>main>div.b>p')
])
expect(groups.map(group => group.label)).toEqual(['div.b', 'div.a'])
expect(groups[0]?.items.map(entry => entry.number)).toEqual([1, 3])
})
it('survives a batch with no element comments at all', () => {
const groups = groupAnnotations([item(1), item(2)])
expect(groups).toHaveLength(1)
expect(groups[0]?.items).toHaveLength(2)
})
})
describe('groupAnnotations refinement', () => {
// A normal page: header / main / footer part company at the top, so a single
// split buries every section under `main`.
const page = [
...['a.logo', 'ul.links>li', 'button.menu'].map((tail, index) => item(index + 1, `body>header.nav>${tail}`)),
...['h1', 'p.sub', 'a.cta', 'img.art'].map((tail, index) => item(index + 4, `body>main>section.hero>${tail}`)),
...['table', 'button.buy', 'span.note'].map((tail, index) => item(index + 8, `body>main>section.pricing>${tail}`)),
...['details:nth-of-type(1)', 'details:nth-of-type(4)'].map((tail, index) =>
item(index + 11, `body>main>section.faq>${tail}`)
),
...['div.cols>ul', 'small.copy'].map((tail, index) => item(index + 13, `body>footer.foot>${tail}`))
]
it('breaks up the branch that would otherwise swallow most of the batch', () => {
const groups = groupAnnotations(page)
const labels = groups.map(group => group.label)
expect(labels).toContain('section.hero')
expect(labels).toContain('section.pricing')
expect(labels).toContain('section.faq')
expect(labels).not.toContain('main')
})
it('leaves no group holding more than a third of the batch', () => {
const groups = groupAnnotations(page)
const ceiling = Math.max(2, Math.ceil(page.length / 3))
for (const group of groups) {
expect(group.items.length).toBeLessThanOrEqual(ceiling)
}
})
it('loses and duplicates nothing while refining', () => {
const numbers = groupAnnotations(page)
.flatMap(group => group.items.map(entry => entry.number))
.sort((a, b) => a - b)
expect(numbers).toEqual(page.map(entry => entry.number))
})
it('keeps refined groups on non-overlapping subtrees', () => {
const keys = groupAnnotations(page).map(group => group.key)
const nested = keys.filter(key => keys.some(other => other !== key && other.startsWith(`${key}>`)))
expect(nested).toEqual([])
})
it('stops instead of looping when an oversized group cannot divide further', () => {
const identical = Array.from({ length: 9 }, (_, index) => item(index + 1, 'body>div.card>span'))
const groups = groupAnnotations(identical)
expect(groups).toHaveLength(1)
expect(groups[0]?.items).toHaveLength(9)
})
})

View File

@@ -0,0 +1,170 @@
/**
* Structural grouping for a comment batch.
*
* Twenty-three comments used to arrive as twenty-three flat blocks, so the
* agent made twenty-three todos and worked them one at a time. The fix is not
* a classifier in the renderer — "is this a UI nit or a functional bug" is a
* judgment only the model can make, and prose-matching it here would be wrong
* constantly. What the renderer CAN know is structure: which pins sit in the
* same part of the DOM, and therefore which ones are likely the same component
* and the same source file.
*
* So this splits the batch by shared ancestor path and hands the model groups
* that touch disjoint subtrees. Disjoint is the property that makes parallel
* work safe — grouping by theme instead ("all the UI ones") would put five
* agents in the same files. The model still owns the semantics and can regroup;
* these are labelled starting points, not orders.
*
* Two properties keep the split honest without a tuning knob:
*
* - It compares ANCESTOR paths, not full selectors. Two comments on the heading
* and the paragraph of one card differ at the leaf, and splitting there would
* hand out singletons — the thing this exists to prevent. Their parents are
* identical, so they group.
* - The depth is derived, then refined: descend the shared prefix until it
* stops being shared, and sub-split any group that ends up holding most of
* the batch. So the group count follows the page the user commented on rather
* than a constant someone picked.
*/
import type { ComposerReadyAnnotation } from './pack'
export interface AnnotateGroup {
/** Shared ancestor prefix, or '' for the group that has no element. */
key: string
items: ComposerReadyAnnotation[]
/** Short human label for the shared region, e.g. `section.hero`. */
label: string
}
const SEP = '>'
function segments(selector: string): string[] {
return selector.split(SEP).filter(Boolean)
}
/**
* The element's container. A one-segment selector is its own container —
* dropping to nothing would collide with the unanchored group's empty key.
*/
function ancestorPath(selector: string): string[] {
const parts = segments(selector)
return parts.length > 1 ? parts.slice(0, -1) : parts
}
function prefixAt(parts: string[], depth: number): string {
return parts.slice(0, depth).join(SEP)
}
/**
* First depth at which the ancestor paths stop agreeing.
*
* Grouping by a prefix of this depth yields the top-level regions the user
* touched. When every path is identical there is no boundary and everything
* belongs to one group.
*/
export function annotateSplitDepth(selectors: readonly string[]): number {
const parts = selectors.map(ancestorPath)
if (parts.length < 2) {
return parts[0]?.length ? 1 : 0
}
const shortest = Math.min(...parts.map(list => list.length))
for (let depth = 1; depth <= shortest; depth++) {
const seen = new Set(parts.map(list => prefixAt(list, depth)))
if (seen.size > 1) {
return depth
}
}
// Every path shares the whole of the shortest one: the shorter paths are
// ancestors of the longer ones, so one segment deeper is where they part.
return parts.some(list => list.length > shortest) ? shortest + 1 : shortest
}
function labelFor(key: string): string {
const parts = segments(key)
return parts[parts.length - 1] || ''
}
function bucket(items: readonly ComposerReadyAnnotation[], depth: number): AnnotateGroup[] {
const byKey = new Map<string, AnnotateGroup>()
for (const item of items) {
const key = prefixAt(ancestorPath(item.identity?.selector || ''), depth)
const group = byKey.get(key)
if (group) {
group.items.push(item)
continue
}
byKey.set(key, { items: [item], key, label: labelFor(key) })
}
return Array.from(byKey.values())
}
/**
* One pass of the split leaves the deepest branch lumped together: on a normal
* page `header`, `main`, and `footer` part company at the top, so every comment
* inside `main` — hero, pricing, faq — lands in one oversized group. That group
* is not foldable into a single change and not safely divisible among workers,
* which is the whole point of grouping.
*
* So refine: while some group holds more than a third of the batch and its
* members do diverge further down, replace it with its own sub-split. A group
* holding most of the batch has not separated anything. Each pass strictly
* shrinks the largest group or finds it indivisible, so this terminates.
*/
function refine(groups: AnnotateGroup[], total: number): AnnotateGroup[] {
const ceiling = Math.max(2, Math.ceil(total / 3))
let current = groups
for (let pass = 0; pass < total; pass++) {
const target = current.find(group => group.items.length > ceiling)
if (!target) {
break
}
const selectors = target.items.map(item => item.identity?.selector || '')
const deeper = annotateSplitDepth(selectors)
const split = bucket(target.items, deeper)
if (split.length < 2) {
break
}
current = current.flatMap(group => (group === target ? split : [group]))
}
return current
}
/**
* Split a packed batch into groups the model can hand out in parallel.
*
* Comments with no element (area pins) cannot be placed in the tree, so they
* collect in one trailing group rather than being guessed into someone else's
* subtree. Group order follows first appearance, so numbering still reads in
* the order the user clicked.
*/
export function groupAnnotations(items: readonly ComposerReadyAnnotation[]): AnnotateGroup[] {
const placed = items.filter(item => item.identity?.selector)
const loose = items.filter(item => !item.identity?.selector)
const depth = annotateSplitDepth(placed.map(item => item.identity?.selector || ''))
const groups = refine(bucket(placed, depth), placed.length)
if (loose.length) {
groups.push({ items: [...loose], key: '', label: '' })
}
return groups
}

View File

@@ -1,4 +1,5 @@
export { type AnnotateFlushPorts, type AnnotateFlushResult, flushAnnotateStack } from './flush'
export { type AnnotateGroup, annotateSplitDepth, groupAnnotations } from './group'
export { compactIdentity, type CompactIdentity, type ElementSnapshot, formatIdentityLine } from './identity'
export {
ANNOTATE_HOST_TAG,

View File

@@ -192,3 +192,79 @@ describe('flushAnnotateStack', () => {
expect(annotateFlushPrompt(stacked)).toContain('2 comments')
})
})
describe('annotateFlushPrompt batching', () => {
function at(number: number, selector: string): AnnotatePin {
return pin({
id: `annotate-${number}`,
number,
identity: { css: {}, html: '', selector, tag: 'div', text: '' }
})
}
const batch = packageAnnotateStack([
at(1, 'body>main>section.hero>h1'),
at(2, 'body>main>section.hero>p'),
at(3, 'body>main>section.pricing>button'),
at(4, 'body>main>section.pricing>span'),
at(5, 'body>main>section.faq>li')
])
it('heads each region so a long batch is fewer pieces of work than comments', () => {
const prompt = annotateFlushPrompt(batch, 'http://localhost:5173/')
expect(prompt).toContain('Group 1 — `section.hero` (2 comments)')
expect(prompt).toContain('Group 2 — `section.pricing` (2 comments)')
expect(prompt).toContain('Group 3 — `section.faq` (1 comment)')
expect(prompt).toContain('Work them as 3 pieces of work, not 5.')
})
it('warns against the theme split that would put workers in the same files', () => {
const prompt = annotateFlushPrompt(batch)
expect(prompt).toContain('delegate whole groups')
expect(prompt).toContain('never form new groups by theme')
expect(prompt).toContain('Regroup if the code disagrees')
})
it('still lists every comment exactly once', () => {
const prompt = annotateFlushPrompt(batch)
for (const item of batch) {
expect(prompt.split(`Comment ${item.number}\n`)).toHaveLength(2)
}
})
it('leaves a short batch flat — grouping two comments is noise', () => {
const prompt = annotateFlushPrompt(batch.slice(0, 2))
expect(prompt).not.toContain('Group 1')
expect(prompt).not.toContain('pieces of work')
})
it('leaves a batch flat when every comment is in one region', () => {
const prompt = annotateFlushPrompt(
packageAnnotateStack([
at(1, 'body>div.card>h1'),
at(2, 'body>div.card>p'),
at(3, 'body>div.card>a'),
at(4, 'body>div.card>span')
])
)
expect(prompt).not.toContain('Group 1')
})
it('gives dragged areas their own section instead of a guessed region', () => {
const prompt = annotateFlushPrompt(
packageAnnotateStack([
at(1, 'body>main>section.hero>h1'),
at(2, 'body>main>section.pricing>button'),
pin({ id: 'annotate-3', identity: undefined, kind: 'area', number: 3 }),
at(4, 'body>main>section.faq>li')
])
)
expect(prompt).toContain('Unanchored (dragged areas) (1 comment)')
})
})

View File

@@ -1,3 +1,4 @@
import { type AnnotateGroup, groupAnnotations } from './group'
import { type CompactIdentity, formatIdentityLine } from './identity'
import type { AnnotatePin } from './stack'
@@ -66,16 +67,73 @@ export function packageAnnotateStack(pins: readonly AnnotatePin[]): ComposerRead
return pins.map(packageAnnotatePin)
}
/** Below this a flat list is easier to read than a set of headed sections. */
const GROUP_THRESHOLD = 4
function groupHeading(group: AnnotateGroup, index: number): string {
const what = group.label ? `\`${group.label}\`` : 'Unanchored (dragged areas)'
return `Group ${index + 1} — ${what} (${group.items.length} comment${group.items.length === 1 ? '' : 's'})`
}
/**
* How to work a batch this size.
*
* Two things the model gets wrong when handed a long flat list: it makes one
* task per comment and grinds through them serially, and — told to parallelize
* — it splits by theme, which puts several workers in the same component. So
* say both. The groups below are structural (disjoint DOM subtrees, so usually
* disjoint files), which is what makes handing them out concurrently safe;
* "all the styling ones" is not.
*
* It stays advice, not instruction: the model can see whether these comments
* are really one refactor, and a grouping computed from selectors cannot.
*/
function batchGuidance(groupCount: number, total: number): string {
return [
`These ${total} comments are pre-grouped by where they sit in the page — each group is a different part of the DOM, so the groups should touch mostly separate files.`,
`Work them as ${groupCount} pieces of work, not ${total}. Fold comments in the same group into one change.`,
'If you delegate, delegate whole groups — never split one group across workers, and never form new groups by theme (all the spacing ones, all the copy ones): those cut across the same files and the workers will collide.',
'Regroup if the code disagrees with this split — it is derived from the page structure, not from your source layout.'
].join(' ')
}
export function annotateFlushPrompt(items: readonly ComposerReadyAnnotation[], pageUrl?: string): string {
const where = pageUrl ? ` on ${pageUrl}` : ''
const count = items.length
const header =
count === 1
? `I left a comment${where} in the in-app browser. Address it and keep the scope narrow.`
: `I left ${count} comments${where} in the in-app browser. Address them and keep the scope narrow.`
if (count === 1) {
return [
`I left a comment${where} in the in-app browser. Address it and keep the scope narrow.`,
'',
...items.map(item => item.prompt)
].join('\n')
}
return [header, '', ...items.map(item => item.prompt)].join('\n')
const groups = groupAnnotations(items)
if (count < GROUP_THRESHOLD || groups.length < 2) {
return [
`I left ${count} comments${where} in the in-app browser. Address them and keep the scope narrow.`,
'',
...items.map(item => item.prompt)
].join('\n')
}
const sections = groups.flatMap((group, index) => [
groupHeading(group, index),
...group.items.map(item => item.prompt),
''
])
return [
`I left ${count} comments${where} in the in-app browser. Address them and keep the scope narrow.`,
batchGuidance(groups.length, count),
'',
...sections
]
.join('\n')
.trimEnd()
}
export function dataUrlToBlob(dataUrl: string): Blob {

View File

@@ -44,7 +44,7 @@ The center of the app. You get:
- **The same conversation history** as every other Hermes surface — sessions started here resume in the CLI/TUI and vice versa.
- **Drag-and-drop files** anywhere in the chat area to attach them to your next message.
- **A right-hand preview rail** — render web pages, files, and tool outputs side by side while you keep chatting.
- **Comment mode in the in-app browser** — click **Annotate** in the preview browser bar, then click any element (or drag a box) on the live page and type a note; each saved comment stays as a numbered pin on the page. Saving a pin never sends a turn — when you're done, **Add N comments** attaches a cropped screenshot per pin and a short prompt naming each comment to the composer, and you still hit send yourself. Each element comment carries its CSS selector, its markup, and the computed styles that matter for layout, so the agent can find the element in your source instead of guessing from the picture. Password and hidden field values, and any attribute that looks like a key or token, are redacted on the page before the markup leaves it. Pin numbers hold steady if you delete one, and switching chats clears the stack.
- **Comment mode in the in-app browser** — click **Annotate** in the preview browser bar, then click any element (or drag a box) on the live page and type a note; each saved comment stays as a numbered pin on the page. Saving a pin never sends a turn — when you're done, **Add N comments** attaches a cropped screenshot per pin and a short prompt naming each comment to the composer, and you still hit send yourself. Each element comment carries its CSS selector, its markup, and the computed styles that matter for layout, so the agent can find the element in your source instead of guessing from the picture. Password and hidden field values, and any attribute that looks like a key or token, are redacted on the page before the markup leaves it. Larger batches arrive grouped by which part of the page each comment sits in, so twenty-odd comments become a handful of pieces of work rather than one task each — and because the groups are separate DOM subtrees they usually touch separate files, which is what makes handing them to parallel workers safe. Pin numbers hold steady if you delete one, and switching chats clears the stack.
- **Composer history and queue editing** — press the up/down arrow keys in an empty composer to recall and reuse previous prompts, and edit messages you've queued up before they're sent. Pressing Stop (or Esc) while turns are queued pauses the queue and expands it above the composer; resume it from there, or send, edit, and delete individual entries.
- **A conversation timeline rail** — long chats get a slim rail of markers along the edge of the transcript, one per prompt. Hover it to pop open the list of prompts, click one to jump straight to that point in the conversation. (It appears once the chat has a handful of turns.)
- **Find in page** — press **Cmd/Ctrl+F** to open a find bar that searches the rendered chat transcript. Enter / Shift+Enter (or Cmd/Ctrl+G / Cmd/Ctrl+Shift+G while the bar is open) step through matches; Esc closes it.