Skip to content

Commit 0090931

Browse files
committed
Add saliency: spectral-residual visual saliency (where to look)
When there's no template, colour or text to key on, an agent still needs a cue for where to look. Compute the spectral-residual saliency map (Hou & Zhang 2007) and rank salient boxes in source coordinates. Pure numpy FFT (cv2.saliency is opencv-contrib, forbidden), reusing visual_match's grayscale loader and cv2_utils.blobs.connected_boxes; regions threshold at mean+2*std by default. A coarse attention cue to narrow where a template / OCR pass then looks.
1 parent 538a6b4 commit 0090931

11 files changed

Lines changed: 369 additions & 0 deletions

File tree

‎WHATS_NEW.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,11 @@
11
# What's New — AutoControl
22

3+
## What's new (2026-06-24) — Visual Saliency (where to look — spectral-residual)
4+
5+
Find the region that stands out, with no template / colour / text. Full reference: [`docs/source/Eng/doc/new_features/v190_features_doc.rst`](docs/source/Eng/doc/new_features/v190_features_doc.rst).
6+
7+
- **`saliency_map` / `salient_regions` / `most_salient`** (`AC_salient_regions`, `AC_most_salient`): when there's no template, colour or text to key on, an agent still needs a cue for *where to look*. This computes the spectral-residual saliency map (Hou & Zhang 2007 — log amplitude minus its local average, reconstructed through the phase) and turns it into ranked salient boxes in source pixel coordinates. The transform is a pure numpy FFT (`cv2.saliency` is in the forbidden opencv-contrib package, so it's re-implemented over base opencv); it reuses `visual_match`'s grayscale loader and `cv2_utils.blobs.connected_boxes`. Regions threshold at `mean + 2·std` by default. A coarse attention cue to *narrow* where a template / OCR pass then looks. No `PySide6`.
8+
39
## What's new (2026-06-24) — Display-Scale / Visual-DPI Detection
410

511
Infer which display scale (DPI) a template renders at — and how confidently. Full reference: [`docs/source/Eng/doc/new_features/v189_features_doc.rst`](docs/source/Eng/doc/new_features/v189_features_doc.rst).
Lines changed: 49 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,49 @@
1+
Visual Saliency (where to look — spectral-residual)
2+
===================================================
3+
4+
When there is no template, no known colour and no text to OCR, an agent still
5+
needs a cue for *where to look* — the region that stands out from its
6+
surroundings (a popup, a badge, a highlighted row). ``saliency`` computes the
7+
spectral-residual saliency map (Hou & Zhang 2007) — ``log`` amplitude minus its
8+
local average, reconstructed through the phase — and turns it into ranked salient
9+
boxes.
10+
11+
* :func:`saliency_map` — the normalised (0–1) saliency map as an ndarray,
12+
* :func:`salient_regions` — ranked salient boxes ``{x, y, width, height, center,
13+
score}`` in source pixel coordinates,
14+
* :func:`most_salient` — the single most salient region (the first place to look).
15+
16+
The transform is a pure ``numpy`` FFT — ``cv2.saliency`` lives in the forbidden
17+
opencv-contrib package, so it is re-implemented over base opencv only. It reuses
18+
``visual_match``'s grayscale loader (any ndarray / path / PIL image, or the live
19+
screen) and ``cv2_utils.blobs.connected_boxes`` for region extraction. cv2 /
20+
numpy are lazily imported. Imports no ``PySide6``.
21+
22+
Headless API
23+
------------
24+
25+
.. code-block:: python
26+
27+
from je_auto_control import saliency_map, salient_regions, most_salient
28+
29+
most_salient("screen.png")
30+
# {"x": 612, "y": 40, "width": 180, "height": 36, "center": [702, 58],
31+
# "score": 0.82}
32+
33+
for region in salient_regions("screen.png"): # most-salient first
34+
...
35+
36+
sal = saliency_map("screen.png") # (64, 64) float32 in 0..1
37+
38+
Regions are thresholded at ``mean + 2·std`` of the saliency map by default (pass
39+
``threshold`` to override), extracted with ``connected_boxes`` and scaled back to
40+
the source's pixel coordinates. ``size`` is the (small) resolution the saliency is
41+
computed at. Saliency is a coarse attention cue, not a precise detector — use it
42+
to *narrow* where a template / OCR pass then looks.
43+
44+
Executor commands
45+
-----------------
46+
47+
``AC_salient_regions`` and ``AC_most_salient`` (``source`` / ``region`` / ``size``
48+
/ ``threshold`` / ``min_area``). They are exposed as read-only ``ac_*`` MCP tools
49+
and as Script Builder commands under **Image**.
Lines changed: 42 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,42 @@
1+
視覺顯著度(該看哪裡——spectral-residual)
2+
==========================================
3+
4+
當沒有模板、沒有已知顏色、也沒有文字可 OCR 時,agent 仍需要一個*該看哪裡*的線索——也就是從
5+
周遭凸顯出來的區域(彈出視窗、徽章、被反白的列)。``saliency`` 計算 spectral-residual 顯著度圖
6+
(Hou & Zhang 2007)——``log`` 振幅減去其區域平均,再透過相位重建——並轉成排序後的顯著方框。
7+
8+
* :func:`saliency_map` ——正規化(0–1)的顯著度圖(ndarray),
9+
* :func:`salient_regions` ——排序後的顯著方框 ``{x, y, width, height, center, score}``
10+
(以來源像素座標表示),
11+
* :func:`most_salient` ——單一最顯著的區域(第一個該看的地方)。
12+
13+
此轉換為純 ``numpy`` FFT——``cv2.saliency`` 位於被禁用的 opencv-contrib 套件,故在 base opencv
14+
上重新實作。它重用 ``visual_match`` 的灰階載入器(任何 ndarray / 路徑 / PIL 影像,或存活螢幕)與
15+
``cv2_utils.blobs.connected_boxes`` 做區域擷取。cv2 / numpy 為延遲匯入。不匯入 ``PySide6``。
16+
17+
無頭 API
18+
--------
19+
20+
.. code-block:: python
21+
22+
from je_auto_control import saliency_map, salient_regions, most_salient
23+
24+
most_salient("screen.png")
25+
# {"x": 612, "y": 40, "width": 180, "height": 36, "center": [702, 58],
26+
# "score": 0.82}
27+
28+
for region in salient_regions("screen.png"): # 最顯著者在前
29+
...
30+
31+
sal = saliency_map("screen.png") # (64, 64) float32,範圍 0..1
32+
33+
區域預設以顯著度圖的 ``mean + 2·std`` 為門檻(可傳 ``threshold`` 覆寫),以 ``connected_boxes``
34+
擷取,並縮放回來源的像素座標。``size`` 是計算顯著度所用的(較小)解析度。顯著度是粗略的注意力
35+
線索,而非精確偵測器——用它來*縮小*接著由模板 / OCR 比對的範圍。
36+
37+
執行器指令
38+
----------
39+
40+
``AC_salient_regions`` 與 ``AC_most_salient``(``source`` / ``region`` / ``size`` /
41+
``threshold`` / ``min_area``)。皆以唯讀 ``ac_*`` MCP 工具及 Script Builder 指令(位於 **Image**
42+
分類下)形式提供。

‎je_auto_control/__init__.py‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -84,6 +84,10 @@
8484
)
8585
# Display-scale / visual-DPI detection (per-scale match profile)
8686
from je_auto_control.utils.scale_detect import detect_scale, scale_sweep
87+
# Spectral-residual visual saliency (where to look — map + salient regions)
88+
from je_auto_control.utils.saliency import (
89+
most_salient, salient_regions, saliency_map,
90+
)
8791
# VLM element locator (headless)
8892
from je_auto_control.utils.vision import (
8993
VLMNotAvailableError, click_by_description, locate_by_description,
@@ -1660,6 +1664,7 @@ def start_autocontrol_gui(*args, **kwargs):
16601664
"plan_file_drop", "drop_files",
16611665
"image_quality", "is_blurry", "quality_gate",
16621666
"detect_scale", "scale_sweep",
1667+
"saliency_map", "salient_regions", "most_salient",
16631668
# VLM locator
16641669
"VLMNotAvailableError", "locate_by_description", "click_by_description",
16651670
"verify_description",

‎je_auto_control/gui/script_builder/command_schema.py‎

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -787,6 +787,24 @@ def _add_image_specs(specs: List[CommandSpec]) -> None:
787787
),
788788
description="Per-scale match-score profile of a template.",
789789
))
790+
saliency_fields = (
791+
FieldSpec("source", FieldType.FILE_PATH, optional=True),
792+
FieldSpec("region", FieldType.STRING, optional=True,
793+
placeholder=_REGION_PLACEHOLDER),
794+
FieldSpec("size", FieldType.INT, optional=True, default=64),
795+
FieldSpec("threshold", FieldType.FLOAT, optional=True),
796+
FieldSpec("min_area", FieldType.INT, optional=True, default=4),
797+
)
798+
specs.append(CommandSpec(
799+
"AC_salient_regions", "Image", "Salient Regions",
800+
fields=saliency_fields,
801+
description="Visually salient regions (spectral-residual; where to look).",
802+
))
803+
specs.append(CommandSpec(
804+
"AC_most_salient", "Image", "Most Salient Region",
805+
fields=saliency_fields,
806+
description="The single most visually salient region of an image/screen.",
807+
))
790808
specs.append(CommandSpec(
791809
"AC_changed_regions", "Image", "Changed Regions (motion)",
792810
fields=(

‎je_auto_control/utils/executor/action_executor.py‎

Lines changed: 23 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -4327,6 +4327,27 @@ def _scale_sweep(template: Any, haystack: Any = None, region: Any = None,
43274327
method=str(method))}
43284328

43294329

4330+
def _salient_regions(source: Any = None, region: Any = None, size: Any = 64,
4331+
threshold: Any = None, min_area: Any = 4) -> Dict[str, Any]:
4332+
"""Adapter: ranked visually-salient regions of an image / the screen."""
4333+
from je_auto_control.utils.saliency import salient_regions
4334+
cut = float(threshold) if threshold not in (None, "") else None
4335+
regions = salient_regions(source, region=_coerce_region(region),
4336+
size=int(size), threshold=cut,
4337+
min_area=int(min_area))
4338+
return {"regions": regions, "count": len(regions)}
4339+
4340+
4341+
def _most_salient(source: Any = None, region: Any = None, size: Any = 64,
4342+
threshold: Any = None, min_area: Any = 4) -> Dict[str, Any]:
4343+
"""Adapter: the single most visually-salient region (where to look)."""
4344+
from je_auto_control.utils.saliency import most_salient
4345+
cut = float(threshold) if threshold not in (None, "") else None
4346+
result = most_salient(source, region=_coerce_region(region),
4347+
size=int(size), threshold=cut, min_area=int(min_area))
4348+
return {"found": result is not None, "region": result}
4349+
4350+
43304351
def _image_histogram(source: Any = None, bins: Any = 32, space: str = "hsv",
43314352
region: Any = None) -> Dict[str, Any]:
43324353
"""Adapter: per-channel colour histogram of an image / the screen."""
@@ -6553,6 +6574,8 @@ def __init__(self):
65536574
"AC_quality_gate": _quality_gate,
65546575
"AC_detect_scale": _detect_scale,
65556576
"AC_scale_sweep": _scale_sweep,
6577+
"AC_salient_regions": _salient_regions,
6578+
"AC_most_salient": _most_salient,
65566579
"AC_image_histogram": _image_histogram,
65576580
"AC_histogram_changed": _histogram_changed,
65586581
"AC_changed_regions": _changed_regions,

‎je_auto_control/utils/mcp_server/tools/_factories.py‎

Lines changed: 29 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -3443,6 +3443,35 @@ def img_histogram_tools() -> List[MCPTool]:
34433443
handler=h.scale_sweep,
34443444
annotations=READ_ONLY,
34453445
),
3446+
MCPTool(
3447+
name="ac_salient_regions",
3448+
description=("Visually salient regions of 'source' (image path; "
3449+
"default screen grab of 'region') via spectral-residual "
3450+
"saliency — where to look with no template/text. Returns "
3451+
"{regions:[{x,y,width,height,center,score}], count}."),
3452+
input_schema=schema({
3453+
"source": {"type": "string"},
3454+
"region": {"type": "array", "items": {"type": "integer"}},
3455+
"size": {"type": "integer"},
3456+
"threshold": {"type": "number"},
3457+
"min_area": {"type": "integer"}}),
3458+
handler=h.salient_regions,
3459+
annotations=READ_ONLY,
3460+
),
3461+
MCPTool(
3462+
name="ac_most_salient",
3463+
description=("The single most visually salient region of 'source' "
3464+
"(default screen): {found, region:{x,y,width,height,"
3465+
"center,score}}. The first place to look."),
3466+
input_schema=schema({
3467+
"source": {"type": "string"},
3468+
"region": {"type": "array", "items": {"type": "integer"}},
3469+
"size": {"type": "integer"},
3470+
"threshold": {"type": "number"},
3471+
"min_area": {"type": "integer"}}),
3472+
handler=h.most_salient,
3473+
annotations=READ_ONLY,
3474+
),
34463475
]
34473476

34483477

‎je_auto_control/utils/mcp_server/tools/_handlers.py‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2532,6 +2532,17 @@ def scale_sweep(template, haystack=None, region=None, scales=None,
25322532
return _scale_sweep(template, haystack, region, scales, method)
25332533

25342534

2535+
def salient_regions(source=None, region=None, size=64, threshold=None,
2536+
min_area=4):
2537+
from je_auto_control.utils.executor.action_executor import _salient_regions
2538+
return _salient_regions(source, region, size, threshold, min_area)
2539+
2540+
2541+
def most_salient(source=None, region=None, size=64, threshold=None, min_area=4):
2542+
from je_auto_control.utils.executor.action_executor import _most_salient
2543+
return _most_salient(source, region, size, threshold, min_area)
2544+
2545+
25352546
def image_histogram(source=None, bins=32, space="hsv", region=None):
25362547
from je_auto_control.utils.executor.action_executor import _image_histogram
25372548
return _image_histogram(source, bins, space, region)
Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,6 @@
1+
"""Spectral-residual visual saliency: map + ranked salient regions (numpy FFT)."""
2+
from je_auto_control.utils.saliency.saliency import (
3+
most_salient, salient_regions, saliency_map,
4+
)
5+
6+
__all__ = ["saliency_map", "salient_regions", "most_salient"]
Lines changed: 101 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,101 @@
1+
"""Find the visually salient regions of a frame (spectral-residual saliency).
2+
3+
When there is no template, no known colour and no text to OCR, an agent still
4+
needs a cue for *where to look* — the region that stands out from its
5+
surroundings (a popup, a badge, a highlighted row). ``saliency`` computes the
6+
spectral-residual saliency map (Hou & Zhang 2007) — ``log`` amplitude minus its
7+
local average, reconstructed through the phase — and turns it into ranked salient
8+
boxes.
9+
10+
The transform is a pure ``numpy`` FFT (``cv2.saliency`` lives in the forbidden
11+
opencv-contrib package, so it is re-implemented here over base opencv only). It
12+
reuses ``visual_match``'s grayscale loader for the source (any ndarray / path /
13+
PIL image, or the live screen) and ``cv2_utils.blobs.connected_boxes`` for the
14+
region extraction. cv2 / numpy are lazily imported. Imports no ``PySide6``.
15+
"""
16+
from typing import Any, Dict, List, Optional, Sequence, Tuple
17+
18+
ImageSource = Any
19+
20+
21+
def _gray(source: Optional[ImageSource], region: Optional[Sequence[int]]):
22+
from je_auto_control.utils.visual_match.visual_match import _haystack_gray
23+
return _haystack_gray(source, region)
24+
25+
26+
def _saliency_from_gray(gray, size: int):
27+
import cv2
28+
import numpy as np
29+
small = cv2.resize(gray, (size, size),
30+
interpolation=cv2.INTER_AREA).astype(np.float32)
31+
fft = np.fft.fft2(small)
32+
log_amplitude = np.log(np.abs(fft) + 1e-8)
33+
residual = log_amplitude - cv2.blur(log_amplitude, (3, 3))
34+
recon = np.fft.ifft2(np.exp(residual + 1j * np.angle(fft)))
35+
smoothed = cv2.GaussianBlur(np.abs(recon) ** 2, (0, 0), sigmaX=3.0)
36+
peak = float(smoothed.max())
37+
if peak > 0:
38+
smoothed = smoothed / peak
39+
return smoothed.astype(np.float32)
40+
41+
42+
def saliency_map(source: Optional[ImageSource] = None, *,
43+
region: Optional[Sequence[int]] = None, size: int = 64):
44+
"""Return the normalised (0–1) spectral-residual saliency map as an ndarray.
45+
46+
The map is computed at ``size`` x ``size`` (the algorithm's native low
47+
resolution); higher = more salient.
48+
"""
49+
return _saliency_from_gray(_gray(source, region), int(size))
50+
51+
52+
def _regions_from_saliency(saliency, orig_shape: Tuple[int, int],
53+
threshold: Optional[float], min_area: int,
54+
size: int) -> List[Dict[str, Any]]:
55+
from je_auto_control.utils.cv2_utils.blobs import connected_boxes
56+
if threshold is not None:
57+
cut = float(threshold)
58+
else: # scale-invariant: regions standing 2 std above the mean saliency
59+
cut = float(saliency.mean()) + 2.0 * float(saliency.std())
60+
mask = (saliency >= cut).astype("uint8") * 255
61+
orig_height, orig_width = int(orig_shape[0]), int(orig_shape[1])
62+
scale_x, scale_y = orig_width / float(size), orig_height / float(size)
63+
regions: List[Dict[str, Any]] = []
64+
for box in connected_boxes(mask, min_area=min_area):
65+
x, y = int(box["x"] * scale_x), int(box["y"] * scale_y)
66+
width = max(1, int(box["width"] * scale_x))
67+
height = max(1, int(box["height"] * scale_y))
68+
patch = saliency[box["y"]:box["y"] + box["height"],
69+
box["x"]:box["x"] + box["width"]]
70+
score = float(patch.mean()) if patch.size else 0.0
71+
regions.append({"x": x, "y": y, "width": width, "height": height,
72+
"center": [x + width // 2, y + height // 2],
73+
"score": score})
74+
regions.sort(key=lambda region: region["score"], reverse=True)
75+
return regions
76+
77+
78+
def salient_regions(source: Optional[ImageSource] = None, *,
79+
region: Optional[Sequence[int]] = None, size: int = 64,
80+
threshold: Optional[float] = None,
81+
min_area: int = 4) -> List[Dict[str, Any]]:
82+
"""Return salient regions as ``[{x, y, width, height, center, score}]``.
83+
84+
Boxes are thresholded from the saliency map (default cut = 3x the mean,
85+
per Hou & Zhang), extracted with ``connected_boxes`` and scaled back to the
86+
source's pixel coordinates, ranked most-salient first.
87+
"""
88+
gray = _gray(source, region)
89+
saliency = _saliency_from_gray(gray, int(size))
90+
return _regions_from_saliency(saliency, gray.shape[:2], threshold,
91+
int(min_area), int(size))
92+
93+
94+
def most_salient(source: Optional[ImageSource] = None, *,
95+
region: Optional[Sequence[int]] = None, size: int = 64,
96+
threshold: Optional[float] = None,
97+
min_area: int = 4) -> Optional[Dict[str, Any]]:
98+
"""Return the single most salient region, or ``None`` if none stand out."""
99+
regions = salient_regions(source, region=region, size=size,
100+
threshold=threshold, min_area=min_area)
101+
return regions[0] if regions else None

0 commit comments

Comments
 (0)