Skip to content

About

"STE persistence" is shorthand for Simplified Technical English persistence — how long a model keeps obeying the STE constraint before it drifts back to ordinary prose.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

STE Retention

This project answers one question. Does a model follow a writing standard better when you give the standard a name?

The project tests ASD-STE100 Simplified Technical English (STE). The project is a research tool. It is not an STE checker. It does not give STE certification.

The question

You can ask a model to use a controlled writing style in two ways. You can name the standard. You can also describe selected rules.

The two methods ask for the same type of output. However, they can have different effects.

A name can work as a key. The model can use the key to find information that it learned about the standard. A description has no such key. The model reads the description as a set of direct instructions.

This project measures the difference between the two methods. It also measures how the difference changes as the conversation becomes longer.

The design

The experiment crosses two factors. One factor is the name of the standard. The other factor is the list of rules. The two factors make four prompt variants.

Variant Names the standard Gives the rules
bare No No
named Yes No
rules No Yes
named_rules Yes Yes

This cross is necessary. A comparison of a short named prompt with a long rule prompt cannot separate the two effects. The difference can come from the name. The difference can also come from the rule detail or prompt length.

The experiment gives three results:

  • The name effect across both levels of rule detail.
  • The rule effect across both levels of naming.
  • The interaction between the name and the rules.

The code calculates these paired contrasts:

  • Rule effect: ((rules - bare) + (named_rules - named)) / 2
  • Name effect: ((named - bare) + (named_rules - rules)) / 2
  • Interaction: named_rules - named - rules + bare

A negative interaction can show that the name and the rules overlap. The interaction does not, by itself, identify the cause of the overlap.

Pairs

One session runs all four variants against one model. All four variants get the same questions in the same sequence. This procedure gives paired observations.

The stored seed and the session identity determine the sequence. A session does not repeat a question. Saved replies rebuild the same conversation when a run resumes.

The leaderboard compares results only within the same run, model, session, depth, protocol version, and scoring version. It rejects an incomplete four-variant cell. Thus, results from different questions or incompatible versions do not enter one pair.

Probes

The full study scores compliance at selected conversation depths. The default depths are turns 1, 6, and 12.

The worker still sends the turns between the probes. These turns build the conversation context. The worker keeps the complete history for each variant. It scores and records only the selected probe turns.

The web preview is different. It scores each requested turn and permits only one to three turns. Use the preview to examine the workflow. Do not combine preview results with research results.

Study size

The full worker uses this logical-generation formula:

models × sessions × 4 variants × deepest probe

The default web study uses one model, six sessions, four variants, and a deepest probe of turn 12. Thus, it has 288 logical generations. Each normally uses one paid provider request but can use four when all three continuation attempts are needed, for a maximum of 1,152 paid generation requests.

An optional judge adds this number of logical judge units:

models × sessions × 4 variants × number of probe depths

The web form does not use a judge or an approved-word file. The command-line worker can use both options. Judge units have the same maximum of four paid provider requests each. The research safety ceiling is enforced against the combined maximum provider-request count, not the smaller logical-unit count.

Adaptive control

The first historical program used batches and repeated statistical tests. It could stop when the result was clear or when the budget was exhausted. It could also extend the probes to turns 20 and 32.

The current ste.research worker does not implement this adaptive procedure. It runs the models, sessions, and depths that the operator specifies. The budget_usd value is stored with the configuration, but it does not stop provider charges.

This difference is deliberate. The current worker does not claim that fixed settings implement the historical Pocock stopping rule.

Before you start

Use CPython 3.14.7 or a later compatible version. Create a virtual environment and install the development requirements.

python -m venv .venv
. .venv/bin/activate
pip install -r requirements-dev.txt

You need an OpenRouter API key for model calls. The web form sends the key in one request. The command-line worker reads the key from its process environment.

export OPENROUTER_API_KEY='sk-or-v1-...'

The application does not save or log the key. The person who supplies the key is responsible for all charges.

How to get the dictionary

The optional vocabulary score needs a list of approved words. The standard contains the dictionary. This project does not supply the standard or the dictionary.

Procedure

  1. Open the ASD-STE100 downloads page.
  2. Follow the ASD instructions to request an official copy.
  3. Read the license terms that come with the copy.
  4. Find Part 2, the dictionary.
  5. Make a text file only if your license terms permit this use.
  6. Put one approved word on each line.
  7. Use lower-case letters.
  8. Do not include entries that the standard does not approve.
  9. Save the file outside the repository.
  10. Pass the path to python -m ste.research with --approved-words.

Do not commit or publish this word list. The command-line worker reads the file at run time. The worker does not copy the list into a snapshot.

Two cautions

The standard permits applicable technical names and technical verbs. A plain dictionary list cannot identify these project-specific terms. Thus, the vocabulary score can mark a permitted term as unapproved.

The standard is a specification, not only a word list. The current mechanical score tests a small set of properties. It does not test all rules in ASD-STE100.

How to run the experiment

Web application

Start the development server.

./devserver.sh

Open the root page and enter an OpenRouter key. Select a model and an experiment type. The page keeps the key in its masked password field for subsequent requests. It does not place the key in browser storage; clear the field or close the page to remove it.

For every interaction, the server validates both the returned message text and the provider's finish reason. Each generation allows 8,192 output tokens per call. After a token-limit finish, it makes as many as three bounded suffix requests and combines their validated text. Success requires an explicit provider stop; if that does not occur, “Try one prompt” labels the combined text incomplete, and preview and research runners do not score or checkpoint it as completed work. Malformed provider metadata and provider failures produce a sanitized error without returning credentials or the provider response body.

The model menu is built from ste/models/catalog.py, the only place that lists model identifiers, menu labels, and token prices. The page fetches it from GET /api/models, and server validation uses the same catalog. Each menu option shows the label, the exact identifier, and for paid models the input and output prices per million tokens. It currently offers five paid models (Gemini 3.8 Flash, GPT-5.6 Luna, Claude Haiku 4.5, Llama 4 Maverick, and DeepSeek V4 Flash 0731) and seven zero-priced :free variants (Nemotron 3 Ultra, Nemotron 3.5 Lightning, Nemotron 3 Super, Inkling, Qwen3.8 27B, Gemma 4 26B A4B, and Gemma 4 31B). These exact OpenRouter model identifiers and prices were verified against the provider catalog on 2026-09-25. Exact identifiers make a study more reproducible than moving latest aliases, but the catalog can change. Verify availability before you start a large study.

OpenRouter applies per-minute and daily request limits to free variants, and an account's privacy settings can block the providers that serve them. A full study makes hundreds of requests, so a free-model study can stop early; resume it later with its run ID.

  • Short preview: This mode uses all four variants. It has at most 12 logical generations and 48 paid provider requests when every generation needs all three continuations.
  • Full research study: This mode uses the selected model and the full ste.research protocol. Six sessions have 288 logical generations and at most 1,152 paid provider requests.

“Try one prompt” and experiments are separate workflows. The free-form prompt is sent only to /api/interact for a single response. Short previews and full research studies do not use that text; they select their experiment prompts from ste/protocol.py::PROMPT_POOL.

The server sends newline-delimited JSON (NDJSON). Each line is one valid JSON object. The server saves each paid response before it reports completion.

A browser, proxy, or server timeout can stop a full study. Cloud Run applies its service request timeout to the complete request: keeping the window open and receiving streamed logical-unit or ETA events does not restart that deadline. The Cloud Run service timeout is deployment configuration, not an application value; inspect it with gcloud run services describe SERVICE --region REGION and update it with gcloud run services update SERVICE --region REGION --timeout 3600. Cloud Run services permit at most 3,600 seconds, so they cannot accept a 21,600-second request.

The 120-second provider_timeout and judge_timeout values in research snapshots instead bound each individual OpenRouter operation. They do not impose a 120-second lifetime on the whole streamed study. If the deployed service has a 120-second request timeout, that external deadline is the likely cause of a stream ending at two minutes even while progress was arriving. Copy the run ID and resume with the same mode and settings after a service timeout.

For an uninterrupted six-hour allowance, deploy the CLI worker as a Cloud Run Job and set its task timeout with gcloud run jobs update JOB --region REGION --task-timeout 21600s. Jobs are the unattended execution path; the browser stream remains the resumable interactive path.

Snapshots created before durable paid-attempt accounting cannot prove how many failed provider sends occurred. They remain readable, but their full request allowance is conservatively treated as consumed, so they cannot make more paid calls. Start a new run instead of resuming such a legacy snapshot.

The web application has no authentication or authorization layer. A person who knows a run ID can inspect status, resume the run, or delete the snapshot. Deletion returns a conflict while that run holds an active lease, which prevents a later checkpoint from recreating the deleted snapshot. Use a suitable deployment boundary if model text is sensitive.

Command-line worker

This command runs two models and uses the optional word list and judge:

python -m ste.research \
  --models google/gemini-3.8-flash openai/gpt-5.6-luna \
  --sessions 6 --depths 1 6 12 --seed 20260916 \
  --budget-usd 40 --provider-timeout 120 \
  --approved-words /secure/approved_words.txt \
  --judge-model anthropic/claude-haiku-4.5 --judge-timeout 120 \
  --state /durable/research/RUN.json --yes

The command shows 576 generation and 144 judge logical units for this example, plus the maximum of 2,880 paid provider requests after bounded continuations. The --yes option confirms this displayed workload. It does not confirm a price estimate.

To resume, use the same state path and all the same configuration values. Add --run-id with the saved run ID. The worker rejects a changed configuration before it makes a new paid call.

Cost

The command-line worker does not contain a pricing table or calculate a currency estimate. The web form uses the dated price snapshot in ste/models/catalog.py only for its explicitly labeled rough planning range; provider prices and model identifiers can change.

The old program estimated 5.92 USD for one default batch. It estimated 12 to 18 USD for a typical adaptive run. These historical values do not describe the current fixed-session worker. Do not use them as a current quote.

Before a run, calculate both logical units and the maximum provider-request count from the formulas in the design section. Check the current prices for each selected model in OpenRouter. Include the growing conversation history, output-token limit, optional judge units, and up to three continuation requests per unit in your estimate.

The CLI --budget-usd option records the operator's approved value. The application does not enforce this value against live provider charges. Set a provider-side spending limit at or below the approved amount.

The web form recalculates the logical-generation count, maximum paid provider-request count, and a rough model-specific cost range when the user changes the model, preview turns, batches, or research sessions. Prices were verified against OpenRouter's live catalog on 2026-09-25. The narrower range is calibrated from completed 288-logical-unit studies that cost $0.08 with Llama 4 Maverick and $2.10 with Gemini 3.8 Flash. It normalizes those observations by each reference model's combined base input/output price, then scales them by the selected model's combined base price and the selected logical workload. This is an empirical planning range, not a quote or spending cap: input/output mix, response length, continuations, and routing can differ. The application caps research at 100 sessions and 20,000 worst-case paid provider requests; the web study reaches 19,200 requests at that session maximum. For a free variant, the form shows a $0.00 estimate with a rate-limit warning instead of a range. The user who enters the OpenRouter key accepts the provider charges and should set a provider-side spending limit.

The files

ste/research.py

This module runs the full experiment. It also supplies the python -m ste.research command.

Configuration. Command options set the models, sessions, probe depths, seed, timeouts, budget value, optional word list, optional judge, state path, and resume ID.

Workload confirmation. The module calculates generation and judge logical-unit counts and their combined maximum paid provider-request count. It prints those counts and the configured budget value. It requires --yes before paid work starts, and rejects a worst-case request count above the research safety ceiling.

One session. The module sends one system instruction for each variant. It keeps all user and assistant messages in that variant's history. It sends the same prompt sequence to all four variants.

Checkpoints. The module saves each generation call and each judge call before it reports the unit. It uses stable unit identities. A resumed run skips saved units and reconstructs the missing conversation history.

Records. The module creates one analytical record at each probe depth. Records contain schema, protocol, and scoring versions. They also contain stable run, session, prompt, model, variant, and depth values.

ste/experiment.py

This module runs the short web preview. It limits synchronous work to one batch, three turns, and 12 units. It stops before the deployment request limit and uses short provider timeouts.

Preview records have run_mode set to preview. The normal leaderboard does not combine them with research data. After a preview finishes, the browser displays each variant's batch, turn, score, and model response while explaining that the saved preview is not published on the research leaderboard.

ste/scoring.py

This module calculates sentence-length and approximate active-voice components. It can add vocabulary and judge components when the caller supplies them.

The composite score is the arithmetic mean of available component percentages. A missing optional component has a null value. It is not a pass or a failure.

The sentence-length component uses a 25-word ceiling. The output also includes the historical diagnostic for sentences of more than 20 words.

ste/statistics.py

This module supplies a dependency-free paired t-test. It calculates the t value, degrees of freedom, two-sided p value, mean effect, effect size, and confidence-interval half-width.

The current research runner does not use this test for adaptive stopping. The module remains available for analysis.

ste/leaderboard.py

This module renders complete research snapshots. It calculates the named-without-rules score plus the rule, marginal naming, and interaction contrasts for each complete paired cell.

The current page is a table sorted by the mean score for the named prompt without spelled-out rules, equivalent to the bare baseline plus the named-minus-bare change. It shows numerical rank and model first, then that named-arm score, bare baseline, three factorial contrasts, paired-session count, total run elapsed time, and a two-sided 95% confidence interval for every estimate, written compactly as mean ± half-width. Only the first metric header says "mean ± 95% CI"; the Understanding uncertainty section defines the notation once. Protocol and scoring versions still isolate incompatible calculations but are retained only as row metadata rather than visible columns. Each effect is calculated inside a complete matched four-arm cell, then repeated depths within the same run and session are averaged into one sampling unit. The baseline interval likewise uses only the matched bare scores. Intervals use the appropriate Student t critical value for the available degrees of freedom; with fewer than two sessions, the point estimate remains visible and the interval shows ± n/a. Elapsed time spans creation through the final write for each completed contributing run, can include pauses during resumed studies, and is unavailable for historical data without valid snapshot timestamps.

The interval calculation treats run/session clusters as independent. It is not a hierarchical analysis, so pooled sessions from different runs may contain variation at more than one level. An interval describes uncertainty in the aggregate estimate; it is not the range of individual scores or a 95% probability statement about the fixed population effect. Wider intervals represent greater uncertainty and should not be reduced to a binary significance label.

The Flask route reads complete research snapshots from EXPERIMENTS_DIR. It can also read the original JSONL format as a migration path.

Web and support files

  • flask-app.py supplies HTTP routes, status checks, fenced leases, and NDJSON streams.
  • index.html is the stand-alone application document.
  • static/app.js supplies browser behavior.
  • static/styles.css supplies presentation.
  • ste/protocol.py supplies prompts, variants, limits, and version values.
  • ste/models/catalog.py is the single list of selectable model identifiers, labels, prices, and cost-calibration anchors. GET /api/models serves it to the browser.
  • ste/runs/ supplies snapshots, backend detection, and conditional object leases.
  • ste/records.py validates record input and the legacy format.
  • originals/ contains historical reference programs. The package build and Ruff exclude these files.

Scoring limits

The score is not a full test of ASD-STE100. Without an approved-word list, the mechanical score measures sentence length and approximate active voice. The optional judge adds a model opinion. The judge is not a certified checker.

Sentence splitting and passive-voice detection use regular expressions. They can give false positive and false negative results. Technical terms can cause false vocabulary failures.

The result applies to the tested model version on the test date. Models can change. The result does not automatically apply to a different model or writing standard.

Per-model results have fewer observations than combined results. Do not use a small per-model result as a general model rank.

The standard

ASD owns ASD-STE100 and its dictionary. Get an official copy from the ASD-STE100 downloads page.

Do not reprint or change the standard. Do not use ASD marks in a way that implies approval. This project does not imply that ASD approves or certifies the software.

The STEMG has published work about ASD-STE100 and artificial intelligence. See the ASD-STE100 and AI white paper.

If you publish experiment results, give the protocol and model date. Publish model replies when your data policy permits this. Do not publish the approved-word file.

Checkpoints, storage, and recovery

The application stores snapshots in EXPERIMENTS_DIR. The default path is /experiments. A research snapshot contains prompts and model replies that are necessary for resume. It never contains the API key.

The service detects the Cloud Storage mount automatically. It coordinates leases and snapshot fencing through conditional object writes. Multiple Cloud Run instances are safe. The web service has no local-filesystem coordination mode.

The web service reads resumed snapshots through the Cloud Storage API. This avoids stale mount-cache data. Snapshot generation preconditions prevent a request that lost its lease from overwriting newer work.

The ste.research CLI with --state remains file-based and takes no lease. Do not point it at a run that the web service might run at the same time.

The leaderboard reads only complete research snapshots. It excludes preview and incomplete research runs. It separates records with different protocol and scoring versions.

The leaderboard opens with a server-rendered SVG dumbbell chart. Each row joins a model's bare-baseline mean (hollow ring) to its named-without-rules mean (filled dot), with thin translucent 95% CI bands. Bands are clamped to the 0 to 100 score axis, and chevrons mark bands that continue past it. The two series colors passed the colorblind-separation checks against the dark surface. Shape also separates the series, native SVG titles give hover readouts, and the table remains the exact, accessible data view. The chart scrolls horizontally on narrow screens rather than shrinking below legibility. The leaderboard presents its results before the uncertainty guidance in a dedicated dark dashboard theme. The experiment form retains its lighter visual treatment, while the leaderboard uses high-contrast cards and a horizontally scrollable comparison table.

Browser behavior and accessibility

During a request, the page shows static/loading.gif in an accessible overlay. The page gives text status. It shows an approximate ETA only after sufficient measured work.

The page has no percentage display or progress bar. It restores controls after success, failure, cancellation, disconnection, or an early stream closure.

The browser inserts external values with textContent. The leaderboard escapes external labels and validates numeric values. The interface supports keyboard use, live status, contrast, and reduced-motion preferences.

Development checks

Run these checks before you commit a change:

pytest
ruff check .
ruff format --check .
python -m compileall -q flask-app.py ste tests

Tests replace paid and external calls with test functions. Tests do not need a provider key, network access, or a Docker daemon.

Runtime packages are in requirements.txt. Development and test packages are in requirements-dev.txt.

Deployment and data care

The Dockerfile is the Cloud Build contract. The container runs Gunicorn as a non-root user. Mount durable storage at /experiments, or set EXPERIMENTS_DIR to a durable path.

Snapshots contain user prompts and model replies. Encrypt storage and backups. Limit operator access. Set and disclose a retention period. Delete expired snapshots and lease files with a scheduled job.

The web service requires a Cloud Storage volume at EXPERIMENTS_DIR in every environment. It refuses lease work without that mount. Application Default Credentials and the service identity's bucket role authorize conditional object operations. No key file is needed.

Do not put keys, request headers, prompts, replies, provider bodies, or licensed words in logs. Logs can contain safe run identifiers, unit identifiers, status values, and timing values.

A full web study can exceed the deployment request limit. Resume after a timeout. For unattended work, run python -m ste.research in a Cloud Run Job.

License

See LICENSE for the software license. ASD-STE100 remains the property of ASD. This project gives no ASD approval or certification.

About

"STE persistence" is shorthand for Simplified Technical English persistence — how long a model keeps obeying the STE constraint before it drifts back to ordinary prose.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages