From 68023291b21af24bda7b620b4fd2643e942fd47e Mon Sep 17 00:00:00 2001 From: AutomatedTester Date: Wed, 12 Aug 2026 12:19:40 +0100 Subject: [PATCH 1/3] [docs] Add a plan for making the published documentation legible to machines MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A growing share of Selenium's documentation is read by models rather than people: during training, during retrieval, and live over HTTP when a coding agent is asked to write a Selenium test. When that reading goes badly the model falls back on stale priors and emits DesiredCapabilities, an executable_path, a third-party driver manager, and a sleep where a wait belongs — and users attribute that code to Selenium. Reviews what this repository publishes today across all five bindings and sets out six tracks of work to fix it, with the waves that can run in parallel. Findings include: the API reference has no machine-readable entry point and is unlinked from the site llms.txt; objects.inv, element-list and xrefmap.yml are generated and then discarded; the Python docs have three competing copies and no canonical URL; deprecation markers exist in every binding but the "gone, use this instead" mapping exists nowhere consumable; and nothing LLM-facing ships inside the wheel, gem, jar, npm package or nupkg. --- docs/plans/llm-facing-documentation.md | 189 +++++++++++++++++++++++++ 1 file changed, 189 insertions(+) create mode 100644 docs/plans/llm-facing-documentation.md diff --git a/docs/plans/llm-facing-documentation.md b/docs/plans/llm-facing-documentation.md new file mode 100644 index 0000000000000..074d728a5860f --- /dev/null +++ b/docs/plans/llm-facing-documentation.md @@ -0,0 +1,189 @@ +# LLM-Facing Documentation + +- Status: Proposed +- Owner: unassigned +- Scope: this repo (API reference + shipped artifacts), with flagged tasks in `SeleniumHQ/seleniumhq.github.io` + +## Purpose + +Selenium's documentation is written for humans reading a browser. A growing share of it is instead +read by models — during training, during retrieval, and live via fetch when a coding agent is asked +to "write a Selenium test". When that reading goes badly, the model falls back on its stale priors +and emits `DesiredCapabilities`, `executable_path=`, `find_element_by_id`, a third-party driver +manager, and a `sleep` where a wait belongs. Users then attribute that code to Selenium. + +This plan makes the documentation this repo produces legible and authoritative to machine readers, +so that the code models generate is the code the project actually recommends. + +## Current state + +### The pipeline + +Five independent generators, one aggregator, one publish step: + +| Binding | Generator | Bazel target | +| --- | --- | --- | +| Java | Javadoc | `//java/src/org/openqa/selenium/grid:all-javadocs` | +| Python | Sphinx + `pydata_sphinx_theme` | `//py:docs` | +| Ruby | YARD | `//rb:docs` | +| JavaScript | JSDoc + `clean-jsdoc-theme` | `//javascript/selenium-webdriver:docs` | +| .NET | DocFX | `//dotnet:docs` | + +`./go all:docs` (`Rakefile:202`) invokes each `:docs` task, which stages HTML into +`build/docs/api/`. `.github/workflows/update-documentation.yml` uploads that as an artifact +and commits it to the `gh-pages` branch under `docs/api/`, served at +`https://www.selenium.dev/selenium/docs/api//`. It is called from `release.yml:258`. + +### Findings + +1. **The API reference is invisible to `llms.txt`.** `https://www.selenium.dev/llms.txt` exists and + is good — it covers the narrative docs and already carries "notes for tools generating Selenium + code". It contains no link to any API reference. `https://www.selenium.dev/selenium/docs/api/llms.txt` + is a 404. The reference is a large, authoritative corpus with no machine-readable entry point. + +2. **HTML only; existing machine-readable indexes are thrown away.** Sphinx emits `objects.inv`, + Javadoc emits `element-list`, DocFX emits `xrefmap.yml`. All three are produced today as + byproducts, none are advertised, linked, or placed at a documented path. + +3. **Docs regenerate only at release.** Every `:docs` task aborts on a nightly/snapshot + version unless passed `force` (`rake_tasks/java.rake:350`, `python.rake:116`, `ruby.rake:133`, + `dotnet.rake:82`, `node.rake:121`), and the workflow only runs from `release.yml`. New API is + undiscoverable to crawlers and fetch-time retrieval for the whole release cycle. + +4. **Python has a duplicate-content problem that actively poisons model output.** Three copies + compete: `selenium.dev/selenium/docs/api/py` (release), `selenium-python-api-docs.readthedocs.io` + (per-commit, per `py/docs/.readthedocs.yaml`), and the unofficial, years-stale + `selenium-python.readthedocs.io` that dominates search and model recall. No `html_baseurl`, no + `rel=canonical` on any of them. `py/docs/source/index.rst` points readers at Read the Docs as the + fresher source, reinforcing the split. + +5. **Fetch fidelity varies sharply by generator.** Javadoc and Sphinx extract cleanly to text. + Fetching the .NET landing page through a normal text extractor failed to return content within a + five-minute budget; DocFX's `outputFormat: apiPage` is commented out in `dotnet/docs/docfx.json` + with the note "generation with errors". YARD and `clean-jsdoc-theme` navigation is client-side. + Where extraction fails, the model gets nothing and substitutes its priors. + +6. **Java has no orientation page.** There is no `overview.html`, so the Javadoc landing page is a + bare table of ~120 package names with no "start here", no task framing, and no link back to the + narrative docs. Per-package prose is in decent shape (112 `package-info.java` across 126 package + directories) — the entry point is the gap. + +7. **Deprecations are not available as data.** Markers exist in source (30 Java files with + `@Deprecated`, 7 JavaScript, 5 Python, 4 Ruby, 3 .NET `[Obsolete]`) and the project has a written + deprecation policy, but there is no single "this is gone, use that instead" dataset. That mapping + is precisely the highest-value artifact for correcting model output, and it does not exist in any + consumable form. + +8. **Nothing LLM-facing ships inside the installed artifact.** The wheel, gem, jar, npm package and + nupkg contain no `llms.txt`, no agent instructions, no pointer to canonical guidance. An agent + working in a user's repo has the package on disk and no grounding from it. + +## Plan + +Six tracks. Track C is five fully independent per-binding tasks. Waves below give the parallelism. + +### Track A — Machine-readable entry points for the API reference + +- **A1. `docs/api/llms.txt`.** Generated as part of `all:docs`. Names each binding's reference, its + stable URL pattern, the location of its symbol inventory (A3), and the guidance dataset (D2). + Mirrors the tone of the existing site `llms.txt`. +- **A2. `docs/api/index.html`.** A real landing page at `/selenium/docs/api/`: pick your language, + what this is versus the narrative docs, current version, link to the previous version. Today that + path has no page of its own. +- **A3. Publish the inventories.** Copy `objects.inv`, `element-list` and `xrefmap.yml` to + documented, stable paths under `docs/api//` and describe their format in A1. Cheap: they are + already generated. +- **A4. Canonical + sitemap.** Emit `rel=canonical` and a per-binding `sitemap.xml`; add the + `Sitemap:` reference for `/selenium/docs/api/` (site repo — see B). + +### Track B — Canonical-URL hygiene (partly cross-repo) + +- **B1. Settle the one canonical home for the Python API docs.** Set `html_baseurl` in + `py/docs/source/conf.py`, have the Read the Docs build emit `rel=canonical` at that base, and + rewrite the `index.rst` wording so the release build is named as authoritative and Read the Docs as + the preview. +- **B2. Link the API reference from the site `llms.txt`.** Cross-repo. Add a per-language API + reference section, and state the relationship: narrative docs for how, API reference for exact + signatures. +- **B3. Reduce recall of the unofficial Python docs.** Cross-repo/outreach. Options, in order of + effort: a clearly-labelled canonical Python API landing page that outranks it, a request to the + maintainer for a deprecation banner and canonical link, and explicit correction in the guidance + dataset (D2). + +### Track C — Make each generator's output extract cleanly (five parallel tasks) + +Each task ends with the same acceptance check: fetch the landing page and one deep symbol page +through a plain text extractor, and assert the class name, method signatures, and parameter docs all +survive. + +- **C1. Java.** Add `overview.html` with orientation and links to selenium.dev; confirm the Javadoc + invocation emits no-frames output and consider `-linksource`. +- **C2. Python.** Add `html_baseurl`, `sphinx.ext.intersphinx`, and `sphinx.ext.linkcode` pointing at + GitHub (`html_show_sourcelink` is on today with nowhere useful to go). +- **C3. Ruby.** Ensure the YARD index is content rather than a frameset and that class pages carry + their own navigation; pass a landing README that orients rather than repeats install steps. +- **C4. JavaScript.** Verify `clean-jsdoc-theme` member pages are server-rendered and readable with + JavaScript disabled; add narrative-doc links to the theme menu in `jsdoc_conf.json`. +- **C5. .NET.** Highest-risk of the five. Reopen `outputFormat: apiPage` in `dotnet/docs/docfx.json`, + or emit a Markdown/plain-text sidecar per type, so the .NET surface is extractable at all. + +### Track D — The guidance dataset + +- **D1. Generate deprecation data from source.** Per binding, extract deprecation markers and their + replacement text into `deprecations.json`; merge into one cross-binding file published under + `docs/api/`. Makes the deprecation policy machine-checkable as a side effect. +- **D2. Curate the anti-pattern → replacement corpus.** The small, high-leverage file: driver paths + and third-party driver managers → Selenium Manager; `DesiredCapabilities` → Options; `sleep` and + mixed implicit/explicit waits → explicit waits; CDP → BiDi; `find_element_by_*` → `By`. Each entry + gets the wrong form, the right form, the version it changed in, and a canonical doc link. Publish + as data, reference from A1 and B2, and reuse the wording already in the site `llms.txt` notes. +- **D3. Publish the ADRs.** `docs/decisions/` holds the authoritative reasoning (for example + `17670-bidi-implementation-boundaries.md`). Include them in the published set and index them in + A1 — "why the API is this way" is what stops a model inventing an alternative. + +### Track E — Ship guidance inside the artifacts + +- **E1.** Add a short `llms.txt` (or `AGENTS.md`) to the wheel, gem, npm package, jar and nupkg: + version, canonical doc URLs, the D2 rules in brief. Every task is independent per binding. +- **E2.** Publish an official Selenium agent skill / MCP-ready description built from D2 and D3, so + agents can be pointed at project-authored guidance instead of inferring it. + +### Track F — Cadence and verification + +- **F1. Nightly docs.** Publish nightly builds to a separate `nightly/` path so new API is + discoverable without clobbering the release reference. The `force` argument already exists. +- **F2. Documentation-coverage gate.** Per-binding CI check that new public API arrives documented. + Small unit-test-shaped checks, not browser tests. +- **F3. Fetch-fidelity check in CI.** Automate the C-track acceptance check across all five + bindings so extraction never silently regresses. +- **F4. Measure it.** A fixed prompt set ("write a Selenium test that…", one per binding) scored for + the D2 anti-patterns. Run once **before** any other work to get a baseline, and again at the end. + Without this the whole plan is unfalsifiable. + +## Parallelism + +``` +Wave 0 (no dependencies — all concurrent) + F4 baseline · C1 · C2 · C3 · C4 · C5 · A3 · D1 · D2 + +Wave 1 (needs Wave 0 output paths / dataset) + A1 · A2 · A4 · B1 · D3 · F1 · F2 · F3 + +Wave 2 (needs canonical URLs from B1, dataset from D) + B2 · B3 · E1 (×5 bindings, concurrent) · E2 + +Wave 3 F4 re-run against the baseline +``` + +Notes on ordering: **F4's baseline must run first** or there is nothing to compare against. D2 is in +Wave 0 because A1, B2 and E1 all consume it; it is editorial work and needs no code. The five C +tasks touch five disjoint directories and five disjoint Bazel targets — no coordination needed +beyond a shared acceptance check. Within Wave 2, the five E1 tasks are likewise disjoint. + +## Risk + +Docs generation touches build/test wiring (`rake_tasks/`, `.github/workflows/`) — tooling with wide +blast radius, per the repo's own guidance. Every track is additive to published output; none changes +public API or wire behavior. C5 (.NET DocFX) carries real risk of breaking the existing .NET docs +build and should land on its own. B2, B3 and A4 require changes in `seleniumhq.github.io` and +coordination with the docs team. From 8923d19f0a8c619d5b1106c3fec6dac1ba266fab Mon Sep 17 00:00:00 2001 From: AutomatedTester Date: Wed, 12 Aug 2026 12:20:08 +0100 Subject: [PATCH 2/3] [py] Publish a machine-readable API reference and ship guidance in the wheel MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The Python API reference is increasingly read by machines — crawlers, retrieval pipelines, and coding agents fetching it live — and it offered them no entry point, no canonical URL, and no way to tell a current API from one removed several releases ago. Three copies of the Python docs compete for recall, none of them declaring which is authoritative. Canonical URL: sets html_baseurl so every page, including the per-commit Read the Docs preview, carries rel=canonical pointing at the released reference. Rewrites the index wording to name that build as canonical and Read the Docs as a preview of the next release. Machine-readable output: a new Sphinx extension writes llms.txt and sitemap.xml into the HTML root; objects.inv is documented at its stable URL; and generate_deprecations.py parses the sources with ast to emit deprecations.json, publishing 16 deprecated APIs with the replacement to use instead. Swaps viewcode for linkcode so source links point at GitHub at the release tag rather than at a second, non-canonical copy of the source rendered into the site. In-package guidance: selenium/llms.txt ships inside the wheel, so a tool with the package on disk finds the project's own guidance instead of inferring it. The published llms.txt reads its rules back out of that file, leaving one copy rather than two to drift apart. Implements the Python tracks of docs/plans/llm-facing-documentation.md. --- .gitignore | 6 + py/BUILD.bazel | 38 ++- py/docs/.readthedocs.yaml | 2 +- py/docs/README.rst | 62 +++++ py/docs/source/_ext/llm_assets.py | 299 ++++++++++++++++++++++++ py/docs/source/_extra/deprecations.json | 136 +++++++++++ py/docs/source/conf.py | 84 ++++++- py/docs/source/index.rst | 13 +- py/generate_deprecations.py | 229 ++++++++++++++++++ py/pyproject.toml | 1 + py/selenium/llms.txt | 72 ++++++ py/test/unit/docs/deprecations_tests.py | 140 +++++++++++ py/test/unit/docs/llm_assets_tests.py | 86 +++++++ rake_tasks/python.rake | 3 +- 14 files changed, 1163 insertions(+), 8 deletions(-) create mode 100644 py/docs/source/_ext/llm_assets.py create mode 100644 py/docs/source/_extra/deprecations.json create mode 100644 py/generate_deprecations.py create mode 100644 py/selenium/llms.txt create mode 100644 py/test/unit/docs/deprecations_tests.py create mode 100644 py/test/unit/docs/llm_assets_tests.py diff --git a/.gitignore b/.gitignore index 33e688dbd08d8..b00b1ba6853dd 100644 --- a/.gitignore +++ b/.gitignore @@ -83,6 +83,12 @@ py/docs/build/ py/docs/source/**/* !py/docs/source/conf.py !py/docs/source/*.rst +# Sphinx extensions, and the deprecation dataset published beside the HTML. +# Sources, unlike the autosummary stubs the rule above is aimed at. +!py/docs/source/_ext/ +!py/docs/source/_ext/*.py +!py/docs/source/_extra/ +!py/docs/source/_extra/*.json py/build/ py/pytestdebug.log py/python.iml diff --git a/py/BUILD.bazel b/py/BUILD.bazel index b4ba64819f2f8..0d3dbc7359afc 100644 --- a/py/BUILD.bazel +++ b/py/BUILD.bazel @@ -541,7 +541,10 @@ py_library( py_library( name = "selenium", srcs = [], - data = ["selenium/py.typed"], + data = [ + "selenium/llms.txt", + "selenium/py.typed", + ], imports = ["."], visibility = ["//visibility:public"], deps = [ @@ -782,6 +785,7 @@ py_test_suite( "--instafail", ], deps = [ + ":docs-tools", ":init-tree", ":selenium", ] + TEST_DEPS, @@ -1395,6 +1399,30 @@ py_binary( deps = [":selenium"], ) +# Parses the workspace sources with `ast` rather than importing them, so it needs +# nothing on the path and is safe to run before anything is built. +py_binary( + name = "generate-deprecations", + srcs = ["generate_deprecations.py"], + main = "generate_deprecations.py", +) + +# The documentation tooling, exposed so its logic can be unit tested rather than +# only exercised by a release-time documentation build. +py_library( + name = "docs-tools", + srcs = [ + "docs/source/_ext/llm_assets.py", + "generate_deprecations.py", + ], + data = ["docs/source/_extra/deprecations.json"], + imports = [ + ".", + "docs/source/_ext", + ], + deps = [":selenium"], +) + py_binary( name = "sphinx-autogen", srcs = ["run_sphinx_autogen.py"], @@ -1417,7 +1445,13 @@ py_console_script_binary( sphinx_docs( name = "docs", - srcs = glob(["docs/source/**/*.rst"]), + srcs = glob([ + "docs/source/**/*.rst", + # Sphinx extensions loaded by conf.py, and files published verbatim + # alongside the HTML via html_extra_path. + "docs/source/_ext/*.py", + "docs/source/_extra/**", + ]), config = "docs/source/conf.py", sphinx = ":sphinx-build", ) diff --git a/py/docs/.readthedocs.yaml b/py/docs/.readthedocs.yaml index f42dd933bb524..68f6bee6c4a40 100644 --- a/py/docs/.readthedocs.yaml +++ b/py/docs/.readthedocs.yaml @@ -13,7 +13,7 @@ build: python: "3.12" commands: - pip install -r py/requirements_lock.txt - - cd py && python3 generate_api_module_listing.py && cd + - cd py && python3 generate_api_module_listing.py && python3 generate_deprecations.py && cd - PYTHONPATH=py sphinx-autogen -o $READTHEDOCS_OUTPUT/html py/docs/source/api.rst - PYTHONPATH=py sphinx-build -b html -d build/docs/doctrees py/docs/source $READTHEDOCS_OUTPUT/html diff --git a/py/docs/README.rst b/py/docs/README.rst index 51de6b80ab922..36a3312421197 100644 --- a/py/docs/README.rst +++ b/py/docs/README.rst @@ -16,3 +16,65 @@ How to build docs After building, docs are available in ``bazel-bin/py/docs/_build/html/`` + + +Canonical URL +============= + +The reference is published from a release at +https://www.selenium.dev/selenium/docs/api/py, and ``html_baseurl`` in +``source/conf.py`` names that address as canonical. Every page carries a +``rel="canonical"`` link back to it, including the per-commit preview built on +Read the Docs, so that the copies do not compete with each other in search +results or in a retrieval index. + +Override the base when serving a build from somewhere else: + +.. code-block:: console + + sphinx-build -b html -D html_baseurl= docs/source _build/html + + +Machine-readable output +======================= + +Much of this reference is read by machines — by crawlers, by retrieval +pipelines, and over HTTP by tools generating Selenium code. Alongside the HTML, +the build publishes four files at stable URLs under +``https://www.selenium.dev/selenium/docs/api/py/``: + +``llms.txt`` + The entry point for machine readers: what the reference covers, where its + indexes are, and the rules for writing correct Selenium code. Generated by + ``source/_ext/llm_assets.py``; the rules themselves are read from + ``py/selenium/llms.txt``, which ships inside the installed package, so + there is only one copy to keep current. + +``objects.inv`` + The Sphinx inventory: every documented symbol mapped to its URL. Written by + Sphinx itself and readable with ``sphinx.ext.intersphinx`` or ``sphobjinv``. + +``deprecations.json`` + Every deprecated API in the release with the replacement to use instead, + extracted from the deprecation warnings in the source by + ``py/generate_deprecations.py``. It is committed under ``source/_extra/`` + and published verbatim through ``html_extra_path``; a unit test fails if it + no longer matches the source. + +``sitemap.xml`` + Every page in the build, so the reference is crawlable without following + client-side navigation. Crawlers only honour ``robots.txt`` at a domain + root, so referencing this sitemap is a change to the + ``seleniumhq.github.io`` repository rather than to this build. + +Regenerate the deprecation dataset on its own with: + +.. code-block:: console + + bazel run //py:generate-deprecations + +The tooling behind these files is unit tested: + +.. code-block:: console + + bazel test //py:unit --test_filter="deprecations or llm_assets" diff --git a/py/docs/source/_ext/llm_assets.py b/py/docs/source/_ext/llm_assets.py new file mode 100644 index 0000000000000..270d6be9aba3e --- /dev/null +++ b/py/docs/source/_ext/llm_assets.py @@ -0,0 +1,299 @@ +# Licensed to the Software Freedom Conservancy (SFC) under one +# or more contributor license agreements. See the NOTICE file +# distributed with this work for additional information +# regarding copyright ownership. The SFC licenses this file +# to you under the Apache License, Version 2.0 (the +# "License"); you may not use this file except in compliance +# with the License. You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, +# software distributed under the License is distributed on an +# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +# KIND, either express or implied. See the License for the +# specific language governing permissions and limitations +# under the License. + + +"""Sphinx extension emitting machine-readable companions to the HTML build. + +A growing share of this reference is read by machines rather than by people: by +crawlers, by retrieval pipelines, and live over HTTP by coding agents asked to +write Selenium code. HTML alone gives those readers no entry point and no way to +tell a current API from one that was removed several releases ago, so they fall +back on stale priors. Two files close that gap, written into the HTML output root +once the build finishes: + +``llms.txt`` + An entry point for machine readers, following the llms.txt convention: what + this corpus is, where its indexes live, and the handful of API-usage rules + that most often come out wrong in generated code. +``sitemap.xml`` + Every page the build produced, so the reference is crawlable without + following client-side navigation. Crawlers only honour ``robots.txt`` at the + domain root, so pointing them at this sitemap is a change to the site + repository, not to this build. + +Both describe absolute URLs, so they are only written when ``html_baseurl`` is +set. Preview builds leave it unset and are skipped rather than advertising URLs +they are not served from. +""" + +from __future__ import annotations + +import os +import posixpath +from xml.sax.saxutils import escape + +# Sphinx is imported inside the hook rather than at module scope so that the +# rendering below stays testable without a documentation toolchain installed. + +# Hand-picked because an alphabetical listing of ~140 modules tells a machine +# reader nothing about where to start. Each entry is a module a generated script +# is likely to need, described in terms of the task it performs. +KEY_MODULES = [ + ( + "selenium_webdriver_remote/selenium.webdriver.remote.webdriver", + "selenium.webdriver.remote.webdriver", + "The WebDriver class every browser driver inherits from: navigation, element lookup, " + "cookies, windows, timeouts, and the driver.script / driver.network / driver.browsing_context " + "BiDi entry points.", + ), + ( + "selenium_webdriver_remote/selenium.webdriver.remote.webelement", + "selenium.webdriver.remote.webelement", + "WebElement: click, send_keys, text, attributes and properties, and nested element lookup.", + ), + ( + "selenium_webdriver_common/selenium.webdriver.common.by", + "selenium.webdriver.common.by", + "The By strategies passed to find_element and find_elements. This is the only supported locator API.", + ), + ( + "selenium_webdriver_support/selenium.webdriver.support.wait", + "selenium.webdriver.support.wait", + "WebDriverWait: the supported way to wait for a condition before acting on it.", + ), + ( + "selenium_webdriver_support/selenium.webdriver.support.expected_conditions", + "selenium.webdriver.support.expected_conditions", + "Ready-made conditions for WebDriverWait, imported by convention as EC.", + ), + ( + "selenium_webdriver_common/selenium.webdriver.common.options", + "selenium.webdriver.common.options", + "The base options class. Browser configuration is passed as options=, never as capability dictionaries.", + ), + ( + "selenium_webdriver_chrome/selenium.webdriver.chrome.options", + "selenium.webdriver.chrome.options", + "Chrome-specific options: arguments, binary location, extensions, preferences.", + ), + ( + "selenium_webdriver_firefox/selenium.webdriver.firefox.options", + "selenium.webdriver.firefox.options", + "Firefox-specific options: arguments, profile, and preferences.", + ), + ( + "selenium_webdriver_common/selenium.webdriver.common.action_chains", + "selenium.webdriver.common.action_chains", + "ActionChains for low-level pointer, key, wheel and pause input.", + ), + ( + "selenium_webdriver_support/selenium.webdriver.support.select", + "selenium.webdriver.support.select", + "Select: working with