enhancement: make preview screenshots fail honestly instead of returning blank images - #54
Merged
Merged
Conversation
Capture reads the window's rendered surface, so it depends on the app having one. Probing a real WebView2 across window states found two failures behind the previous "returns PNG bytes" check: a minimized window makes the browser never answer at all, and a capture taken before the page paints returns valid bytes showing nothing. Capture is now bounded at six seconds and reports that the window needs restoring, instead of consuming the whole operation deadline and surfacing as a vague timeout. Hidden, transparent, and occluded windows were measured and capture normally, so they are left alone. A nearly uniform capture cannot be told apart from a genuinely blank page, so it is flagged rather than refused and the model is told to say the image looks blank instead of describing detail it cannot see. The threshold sits an order of magnitude below a measured painted page. Capturing from the renderer instead of the surface was the obvious candidate fix and is wrong: measured, it throws when minimized and returns a blank image when the window is hidden, which is worse than what it replaces. The smoke test now drives the window through those states rather than asserting only that some bytes came back, and pins the blank threshold against a real painted capture. It also caught that reading a viewport number threw depending on how the value was boxed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #53. The screenshot check there proved only that image bytes came back, not that those bytes showed anything. Probing a real browser across window states found two ways they might not.
What was wrong
A minimized window stops the capture entirely. Not an error — the browser simply never answers. The agent would sit through the full operation deadline and report a vague timeout, with no indication that restoring the window would fix it.
A capture taken before the page finishes drawing returns a valid but empty image. This is the more serious one, because it fails silently: the model receives a picture of nothing and either describes nothing or invents something, having been told it was looking at the page. That runs directly against the principle the rest of this feature follows, which is to refuse rather than fake it.
Neither was visible before, because the old check only asked whether some bytes came back.
What changed
Capture is now bounded at six seconds and, if the browser does not answer, says the window is probably minimized and needs restoring. Hidden, transparent, and covered windows were measured and capture perfectly well, so they are left alone.
An almost-empty capture is flagged rather than rejected. It genuinely cannot be told apart from a real screenshot of a blank page, so the honest response is to hand it over with a warning: the model is told to say the image looks blank rather than describe detail it cannot see. The threshold was set from measurement, an order of magnitude below a real drawn page.
A fix that was measured and rejected
Reading from the browser's renderer rather than the window surface is the obvious candidate and is what the documentation suggests for hidden windows. Measured against a real browser it is worse: it throws when the window is minimized, and returns an empty image when the window is hidden, which is the exact failure it was supposed to prevent. Recorded here so it does not get tried again.
Verification
280 automated tests, up from 279. The browser integration test now drives the window through visible, minimized, and restored states instead of asserting that some bytes came back, and pins the blank threshold against a real drawn capture rather than a guess. It also caught a reading bug where a window measurement threw depending on how the value happened to be stored.
Still not covered
One case needs the real app rather than a test harness: an agent capturing while its tab is in the background. It is plausible that works, since hidden windows capture fine, but it has not been measured and is not claimed.