Skip to content

agent: CMD_NAND, SPI NAND study ops on the FMC100 page engine - #146

Merged
widgetii merged 6 commits into
masterfrom
agent/nand-study
Oct 3, 2026
Merged

widgetii merged 6 commits into
masterfrom
agent/nand-study

Conversation

@widgetii

@widgetii widgetii commented Oct 2, 2026

Copy link
Copy Markdown
Member

Phase 1 of the NAND ECC study for OpenIPC/firmware#2285 (hi3516ev300 + W25N01GV) and #2519 (GK7205V5x0 + GD5F1GM7). It adds tooling to drive the FMC100 SPI NAND path below the OS, both the way the Linux drivers do and deliberately differently.

Agent (protocol v5, CAP_NAND)

  • CMD_NAND 0x0E sub-ops: INFO, FEATURE_GET/SET, READ_PAGES, PROGRAM_PAGE, ERASE_BLOCK, FMC_REG.
  • Two transfer modes:
    • REG: register ops, page + full OOB, raw.
    • DMA: the controller's OP_CTRL page engine, ported from hifmc100_send_cmd_read/write, with an explicit FMC_CFG so the caller picks the ECC type.
  • Each page read records ECC_ERR_NUM0_BUF0, the chip status and FMC_INT.
  • Pages go through a 1 MiB uncached RAM buffer and per-page records through a second buffer. The host moves both with the existing CMD_READ/CMD_WRITE. Running READ_PAGES without storing data is a fast ECC scan (1024 pages in under 2 s).
  • Fixes found on the way:
    • flash_crc32() sent NOR opcodes to NAND, which broke CMD_CRC32 and read-verify on every SPI NAND.
    • The flash-write readback used the FLASH_MEM window, which doesn't exist on NAND and wraps at 1 MiB on NOR.
    • Register-window CRCs used byte loads where reads use word loads.

Host: defib.agent.nand plus defib agent nand info|feature|fmc-reg|read-page|ecc-scan|dump|write-page|erase-block. write-page and erase-block are destructive and listed in CLAUDE.md's safety table. dump is raw by default, with on-die ECC off for its duration.

Verified on hardware

  • INFO and raw paths. W25N01GV reports 2048+64 and GD5F1GM7 2048+128. REG reads and DMA reads with ECC off are byte-identical.
  • On-die ECC at power-on. Both chips come up with it enabled (B0 = 0x18 / 0x10).
  • The #2519 signature on vendor firmware. On the V510's vendor UBI root, scanned with the vendor's 24-bit ECC, exactly 4 pages are uncorrectable. All four read 0x0000FF00, i.e. only step 1 is uncorrectable. They sit at pages 3–4 of PEBs 51 and 52, which is consistent with UBIFS master LEBs (still to be confirmed).
  • On-die ECC flips bits. With the chip's on-die ECC on, 45 pages need correction; with it off, 25.

Checks: pytest 975 passed, ruff, mypy, make -C agent test, clean builds for ev300 / gk7205v500 / hi3520dv200 (HISFC350, feature compiled out) / hi3516cv300 (ARMv5) / cv610.

There's no C-level simulator of the page engine; it would only restate the port. Host tests use a simulated agent that answers CMD_NAND frames.

Groundwork for OpenIPC/firmware#2285 (hi3516ev300 + W25N01GV) and #2519
(GK7205V5x0 + GD5F1GM7): the rootfs rots with uncorrectable ECC that
Linux never reports, and the question is which writes leave pages with
unusable parity. Answering it needs the controller driven below the OS,
the way the Linux hifmc100/xmedia_fmc100 drivers drive it, one knob at a
time.

Agent (v5, CAP_NAND bit 8, FMC100 builds only):
- CMD_NAND 0x0E with sub-ops INFO, FEATURE_GET/SET, READ_PAGES,
  PROGRAM_PAGE, ERASE_BLOCK and FMC_REG (any FMC register, read/write).
- Page I/O in two modes: REG (PAGE_READ/READ_FROM_CACHE or
  PROGRAM_LOAD/EXECUTE through register ops, page + full OOB, no
  controller ECC — raw array with on-die ECC off) and DMA (the
  controller's own OP_CTRL page engine, ported from
  hifmc100_send_cmd_read/write, with an explicit FMC_CFG so ECC type is
  the caller's). Each read records ECC_ERR_NUM0_BUF0, the chip status and
  FMC_INT.
- Pages DMA into a 1 MiB RAM buffer mapped uncached in startup.S (no
  cache maintenance needed); per-page records go to a second buffer. The
  host moves both with the existing CMD_READ/CMD_WRITE stream, so
  CMD_NAND only carries parameters. READ_PAGES without storing data is a
  full-speed ECC scan (1024 pages in under 2 s).
- The NAND chip table now carries OOB size (64 W25N/MX35LF, 128 GD5F).
- Fixes found on the way: flash_crc32() sent NOR read opcodes to a NAND,
  so CMD_CRC32 and `agent read` verify were wrong on every SPI NAND; the
  flash-write readback verified through the FLASH_MEM window, which NAND
  does not have and which wraps at 1 MiB on NOR; CMD_CRC32 over FMC/CRG
  registers used byte loads where CMD_READ uses word loads, so verifying
  a register dump always "mismatched".

Host: defib.agent.nand (NandStudy, records, FMC_CFG composition, the OOB
a kernel write sends) and `defib agent nand info|feature|fmc-reg|
read-page|ecc-scan|dump|write-page|erase-block` (the last two are
destructive, like every writing command). `dump` is raw by default and
turns the chip's on-die ECC off for the duration.

Verified on hardware:
- hi3516ev300 + W25N01GV and GK7205V510 + GD5F1GM7: INFO reports 2048+64
  and 2048+128; REG reads and DMA reads with ECC off return identical
  page + OOB bytes (same MD5), i.e. both are raw.
- Both chips power up with on-die ECC enabled (B0 = 0x18 / 0x10).
- First result on the V510 running its vendor firmware: an ECC scan of
  the vendor UBI root with its 24-bit ECC finds four uncorrectable pages,
  all 0x0000FF00 (step 1 only — the #2519 signature), on pages 3-4 of
  PEBs 51 and 52; with on-die ECC on, 45 pages need correction, with it
  off 25 — the chip's own ECC flips bits when fed foreign parity.

Host tests use a simulated agent that answers CMD_NAND frames as
handle_nand() does. No C-level simulator of the FMC page engine: it
would only restate the port, and the hardware runs above cover it.
@qodo-free-for-open-source-projects

Copy link
Copy Markdown

PR Summary by Qodo

Add FMC100 SPI NAND page-engine study operations

✨ Enhancement 🐞 Bug fix 🧪 Tests 📝 Documentation 🕐 40+ Minutes

Grey Divider

AI Description

• Adds bare-metal raw and controller-ECC NAND operations to investigate page and OOB failures below
 Linux.
• Adds host commands for page reads, ECC scans, dumps, feature inspection, programming, and erasure.
• Fixes NAND CRC and flash-write verification paths; adds simulated host tests and safety guidance.
Diagram

graph TD
  CLI["NAND CLI"] --> Host["NAND study client"] --> Agent["CMD_NAND handler"] --> FMC["FMC100 page engine"] --> Chip[("SPI NAND")]
  Agent --> RAM[("Agent RAM buffers")]
  FMC --> RAM
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Stream page data directly in CMD_NAND
  • ➕ Would avoid separate memory-transfer requests for stored reads and programs.
  • ➖ Would require new bulk-transfer and flow-control behavior despite the existing payload limit.
  • ➖ Would couple NAND study operations more tightly to UART transport handling.

Recommendation: Keep the PR's parameter-only CMD_NAND protocol and reuse CMD_READ/CMD_WRITE for bulk data. It supports record-only ECC scans without transferring pages and avoids duplicating the established streaming path.

Files changed (14) +1529 / -28

Enhancement (9) +1127 / -23
main.cDispatch CMD_NAND and repair flash verification +190/-6

Dispatch CMD_NAND and repair flash verification

• Adds protocol-v5 capability reporting and handlers for NAND information, features, page operations, erasure, and FMC registers. Changes flash-write verification to read through the controller and makes register-window CRC reads consistent with word-based memory reads.

agent/main.c

protocol.hDefine NAND command, response, and status codes +17/-0

Define NAND command, response, and status codes

• Adds CMD_NAND, RSP_NAND, seven sub-operation identifiers, and response statuses to the agent wire protocol.

agent/protocol.h

spi_flash.cDrive FMC100 NAND register and DMA page paths +179/-15

Drive FMC100 NAND register and DMA page paths

• Records chip-specific OOB geometry and adds page reads, programs, erasure, feature access, and FMC100 DMA operations with ECC and status records. Fixes NAND CRC calculation to use NAND reads instead of NOR commands.

agent/spi_flash.c

spi_flash.hExpose NAND geometry, transfers, and page records +56/-2

Expose NAND geometry, transfers, and page records

• Defines register and DMA transfer modes, per-page hardware results, chip geometry, and the NAND study driver interface.

agent/spi_flash.h

client.pyAdvertise the host-side NAND capability bit +1/-0

Advertise the host-side NAND capability bit

• Adds CAP_NAND so host commands can reject agents that lack CMD_NAND support.

src/defib/agent/client.py

nand.pyImplement the host NAND study API +323/-0

Implement the host NAND study API

• Adds transfer configuration, geometry and record decoding, ECC helpers, feature and register access, page operations, and chunked scans. Uses existing agent memory commands to move page data and status records.

src/defib/agent/nand.py

protocol.pyMirror NAND wire codes in Python +4/-0

Mirror NAND wire codes in Python

• Documents and defines CMD_NAND and RSP_NAND alongside the existing host protocol constants.

src/defib/agent/protocol.py

app.pyRegister the agent NAND command group +4/-0

Register the agent NAND command group

• Mounts the new NAND Typer sub-application beneath 'defib agent'.

src/defib/cli/app.py

nand.pyAdd NAND inspection and destructive study commands +353/-0

Add NAND inspection and destructive study commands

• Provides info, feature, FMC-register, page-read, ECC-scan, dump, page-write, and block-erase commands. Dumps default to raw register reads with on-die ECC temporarily disabled and write a JSON record sidecar.

src/defib/cli/nand.py

Tests (1) +340 / -0
test_agent_nand.pyTest NAND host protocol and CLI workflows +340/-0

Test NAND host protocol and CLI workflows

• Adds a simulated CMD_NAND agent and tests transfer packing, ECC decoding, buffered reads, scan chunking and timeouts, programming, erasure, and CLI scan and dump behavior. These tests do not simulate the C FMC100 page engine.

tests/test_agent_nand.py

Documentation (1) +7 / -5
CLAUDE.mdDocument NAND commands and flash-write safety +7/-5

Document NAND commands and flash-write safety

• Adds CMD_NAND and the host subcommands to the project guide. Classifies page programming and block erasure as destructive, while noting that feature and controller-register commands can change volatile state.

CLAUDE.md

Other (3) +55 / -0
MakefileEnable NAND study support on FMC100 builds +2/-0

Enable NAND study support on FMC100 builds

• Defines HAVE_NAND_STUDY only when the FMC100 flash driver is selected, leaving other controller builds without the new command.

agent/Makefile

nand_layout.hDefine shared NAND data and status buffers +28/-0

Define shared NAND data and status buffers

• Declares fixed RAM addresses and sizes for the page/OOB DMA buffer and per-page status records, shared by C and startup assembly.

agent/nand_layout.h

startup.SMap the NAND DMA buffer without caching +25/-0

Map the NAND DMA buffer without caching

• Adds conditional uncached section mappings for the NAND DMA buffer on both ARMv5 and ARMv7, keeping CPU and controller accesses coherent without cache maintenance.

agent/startup.S

@qodo-free-for-open-source-projects

qodo-free-for-open-source-projects Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Incomplete dumps look like successful backups ✓ Resolved
Description
nand_dump() breaks on a short batch but still writes a sidecar and returns normally, while
read_pages() includes the failed page in its returned count and data. When a page read times out,
the output can contain that page's incomplete data and the command exits successfully despite not
backing up the requested range.
Code

src/defib/cli/nand.py[R284-286]

+                    if len(recs) < k:
+                        print(f"\n[controller timeout at page {page}]", file=sys.stderr)
+                        break
Evidence
The agent increments done and stores a record even for the read that fails; the host writes all
returned bytes before checking for a short batch, then finishes normally.

agent/main.c[1314-1324]
src/defib/agent/nand.py[272-284]
src/defib/cli/nand.py[275-299]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A timed-out read produces a partial dump containing the failed page, yet the dump command exits successfully.
## Fix Focus Areas
- src/defib/cli/nand.py[275-299]
- src/defib/agent/nand.py[263-284]
## Recommended Fix
Check each batch's completion and final record before committing its bytes. Preserve the partial output if useful, but mark it incomplete and exit with an error when the requested range was not backed up.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Raw dumps can retain on-die correction ✓ Resolved
Description
nand_dump() discards the readback from feature_set() when disabling on-die ECC and proceeds as
though the requested value took effect. If the chip does not accept that feature write, the default
register-mode dump contains ECC-corrected bytes even though its sidecar says on-die ECC was off.
Code

src/defib/cli/nand.py[R270-271]

+        if not keep_ondie and b0 & 0x10:
+            await study.feature_set(FEATURE_CONFIG, b0 & ~0x10)
Evidence
The agent reads the feature back after writing it and the host exposes that value, but the dump does
not check it before labeling the output raw.

agent/main.c[1296-1301]
src/defib/agent/nand.py[254-256]
src/defib/cli/nand.py[268-296]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The raw-dump command ignores the feature register's readback when switching off on-die ECC.
## Fix Focus Areas
- src/defib/cli/nand.py[268-274]
- src/defib/agent/nand.py[254-256]
## Recommended Fix
Compare the returned configuration with the requested ECC-disabled value before reading pages. Abort the raw dump if ECC remains enabled, while retaining restoration of the original configuration.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Register reads can hide controller stalls ✓ Resolved
Description
nand_page_read() treats the REG helper's returned chip status as proof that its FMC register
operations completed, although fmc_wait_ready() silently stops polling after its timeout. If a
PAGE_READ or READ_FROM_CACHE register operation stalls while a later status read appears valid, the
new page record can report success with stale or incomplete bytes instead of a timeout.
Code

agent/spi_flash.c[R958-962]

+    if (x->mode == NAND_XFER_REG) {
+        uint8_t status = nand_read(page, 0, dst, total);
+        rec->ecc_err = 0;
+        nand_rec_finish(rec, status, 0);
+        return (rec->flags & NAND_REC_TIMEOUT) ? -1 : 0;
Evidence
The wait has a finite counter but no result; the REG helper calls it for the page and cache
commands, and the new study wrapper bases success only on the chip-status timeout flag.

agent/spi_flash.c[220-224]
agent/spi_flash.c[604-640]
agent/spi_flash.c[958-962]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
REG-mode study reads cannot distinguish an FMC register-operation timeout from completion.
## Fix Focus Areas
- agent/spi_flash.c[220-224]
- agent/spi_flash.c[604-640]
- agent/spi_flash.c[958-962]
## Recommended Fix
Make the register-operation wait return completion status, propagate failures from each command in the REG read, and set `NAND_REC_TIMEOUT` rather than returning successful page data.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


4. Per-page interrupt status never reflects the page op ✓ Resolved
Description
nand_rec_finish() reads FMC_INT after nand_get_feature() or nand_wait_oip() has cleared the
interrupt register and started a new register operation. DMA read and program records therefore
capture the status operation’s interrupt bits rather than the page engine’s, which also affects the
fmc_int values in ECC-scan JSON, dump sidecars, and read-page output.
Code

agent/spi_flash.c[R979-981]

+    int timed_out = fmc_wait_dma_done();
+    rec->ecc_err = fmc_reg(FMC_ECC_ERR_NUM0_BUF0);
+    nand_rec_finish(rec, nand_get_feature(NAND_FEATURE_STATUS), timed_out);
Evidence
nand_get_feature() writes 0xFF to FMC_INT_CLR and runs a register operation. In
nand_page_read(), that call is evaluated before nand_rec_finish() reads FMC_INT; the finalizer
therefore sees the subsequent status operation’s interrupt bits rather than the DMA page operation’s
bits. The program path likewise waits through nand_wait_oip() before the finalizer samples the
register.

agent/spi_flash.c[393-402]
agent/spi_flash.c[943-949]
agent/spi_flash.c[979-982]
agent/spi_flash.c[390-405]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`nand_rec_finish()` samples `FMC_INT` after status operations have cleared the page operation’s interrupt bits and started another register operation, so DMA page records contain the later operation’s bits.
## Fix Focus Areas
- agent/spi_flash.c[943-949]
- agent/spi_flash.c[979-982]
- agent/spi_flash.c[1010-1016]
## Recommended Fix
Capture `uint8_t irq = fmc_reg(FMC_INT);` immediately after `fmc_wait_dma_done()` in the read path, before calling `nand_get_feature()`. Capture it in the program path before `nand_wait_oip()`. Pass the saved value to `nand_rec_finish()` and store it in the record instead of reading `FMC_INT` inside the finalizer.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View medium (1)
5. Register CRC still mismatches reads in some cases ✓ Resolved
Description
handle_crc32_cmd() hashes from byte zero of each aligned word instead of the requested address’s
byte offset, and its I/O test omits the per-SoC FMC/CRG/UART windows handled by handle_read.
Unaligned register requests therefore include bytes preceding the requested range, while requests in
omitted windows, such as hi3520dv200 CRG at 0x20030000, still use byte loads; both can make host
read-verification disagree with CMD_READ.
Code

agent/main.c[R303-311]

+    } else if (addr >= 0x10000000 && addr < 0x13000000) {
+        /* Peripheral registers: 32-bit reads only, as CMD_READ does them,
+         * or the CRC covers different bytes than a read returns. */
+        c = 0;
+        for (uint32_t off = 0; off < size; off += 4) {
+            uint32_t val = *(volatile uint32_t *)((addr + off) & ~3u);
+            uint8_t b[4] = { val & 0xff, (val >> 8) & 0xff,
+                             (val >> 16) & 0xff, val >> 24 };
+            c = crc32(c, b, size - off < 4 ? size - off : 4);
Evidence
handle_read computes the byte offset within each word and treats FMC_BASE, CRG_BASE, and UART_BASE
as I/O regions. The new CRC loop instead aligns each word load and hashes starting at b[0],
without applying that offset or the same I/O-region condition, so its byte sequence can differ from
the read path.

agent/main.c[209-253]
agent/main.c[303-312]
agent/main.c[235-248]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`handle_crc32_cmd()` can hash bytes outside the requested range for unaligned register addresses and does not use the same per-SoC I/O windows as `handle_read`, causing CRC results to differ from `CMD_READ`.
## Fix Focus Areas
- agent/main.c[303-312]
- agent/main.c[209-253]
- agent/main.c[235-248]
## Recommended Fix
Use the same `io_region` condition as `handle_read`. For each word, compute `byte_off = (addr + off) & 3`, extract bytes with `(val >> ((byte_off + j) * 8)) & 0xff` only up to the word boundary or end of the requested range, then advance `off` and continue into the next word as needed.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Informational

6. A scan silently changes later raw dumps ✓ Resolved
Description
nand_page_read/nand_page_program write x->fmc_cfg into FMC_CFG and never restore it, while on
NAND the agent sets FMC_CFG only once in flash_init. After any DMA-mode read-page, ecc-scan or
write-page, later CMD_READ, CRC32, feature access and REG-mode dumps (documented as leaving FMC_CFG
untouched) run with the caller's flash select and ECC type.
Code

agent/spi_flash.c[956]

+    if (x->fmc_cfg) fmc_reg(FMC_CFG) = x->fmc_cfg;
Evidence
FMC_CFG is written only in fmc_enter_normal (called from flash_init for NAND) and in the new study
functions, with no restore anywhere.

agent/spi_flash.c[163-195]
agent/spi_flash.c[952-957]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
FMC_CFG set for a DMA op persists for the agent session, which affects later raw dumps and ordinary reads.
## Fix Focus Areas
- agent/spi_flash.c[952-983]
- agent/spi_flash.c[988-1019]
## Recommended Fix
Save fmc_reg(FMC_CFG) before writing x->fmc_cfg and restore it on every return path. Alternatively, save and restore it once per CMD_NAND request in handle_nand.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Tip of the day
💡 Did you know, you can copy the agent prompt from any finding and feed it to your IDE agent

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread src/defib/cli/nand.py Outdated
Comment thread src/defib/cli/nand.py Outdated
Comment thread agent/spi_flash.c
Comment thread agent/spi_flash.c Outdated
Comment thread agent/main.c Outdated
Comment thread agent/spi_flash.c
…y failed

Review (Qodo) of CMD_NAND:

- A dump that hit a controller timeout wrote the failed page's stale
  buffer bytes, a sidecar, and exited 0. It now keeps only pages that
  read, marks the sidecar complete=false with the failed page, and
  exits 1.
- The raw dump assumed the on-die ECC switched off; it now checks the
  feature read-back and refuses rather than label corrected data raw.
- REG-mode reads could not tell a stalled FMC register op from success
  (fmc_wait_ready() just stopped polling). It now reports the timeout,
  and the page record carries NAND_REC_TIMEOUT.
- Records sampled FMC_INT after the following status read had already
  cleared it and run its own op; FMC_INT is now captured right after the
  page op.
- A DMA op's FMC_CFG outlived the request and changed how later
  CMD_READ, CRC32 and raw reads ran; handle_nand restores it after every
  page op (FMC_REG still writes on purpose).
- CMD_CRC32 over registers now uses exactly CMD_READ's I/O windows and
  per-byte word extraction, so unaligned ranges and the V1 FMC/CRG/UART
  windows verify too.
A 140 MB raw dump is ~265 one-megabyte reads over the UART. One sequence
gap made read_memory raise while the agent kept streaming the rest of
that megabyte; the dump's cleanup then sent FEATURE_SET, read a stale
RSP_DATA as its answer, and that error replaced the real one — on
hi3516ev300 after 23k of 65k pages.

CMD_NAND calls now skip the tail of an abandoned read (RSP_DATA frames
and the ACK that closes them; an ACK with no data before it is still a
real answer), dump batches are retried up to three times after waiting
for the line to go quiet, and restoring the on-die ECC setting can no
longer mask the dump's own error.
With the SPI NAND interface selected in FMC_CFG, a non-zero ECC type also
applies to register-path READ_FROM_CACHE data: the controller "corrects"
it, and with an ECC type the page was not written with it flips bits that
were right. Measured on a GK7205V510 + GD5F1GM7: register reads under
FMC_CFG 0x1863 (24-bit) and 0x1823 (8-bit) differ from reads under 0x1803
(ECC off) by 1-6 bits on ~29% of pages, deterministically, while DMA reads
with ECC off match the ECC-off register reads exactly. A raw dump taken
after a DMA scan had left 0x1863 behind was therefore not raw.

REG-mode page reads and programs now clear the ECC type for the op;
handle_nand restores the caller's FMC_CFG afterwards as before.
Fault-injected on a hi3516ev300 (agent at 921600, host forced to 115200)
after real dumps on two boards died the same way:

- _recv_packet_sync() relied on the port's read timeout to recheck its
  deadline, but SerialTransport opens with timeout=None, so read(1)
  blocked until a byte arrived: with the agent silent (wrong baud,
  crashed), every agent command on a plain serial port waited forever,
  whatever timeout it was given. It now polls a blocking port. The
  existing timeout test missed it because its fake port never blocked;
  a new one does.
- recv_response() restarts its timeout on every READY, so a request
  lost to a baud mismatch kept CMD_NAND waiting forever once the agent
  idled back to 115200 and started announcing itself. _call() now uses
  recv_packet() against its own deadline and skips READY itself, as
  well as any other leftover frame (a late answer to an earlier
  request, the tail of an abandoned read).
- NandStudy.resync() treats READY as quiet, re-probes the agent at the
  current rate, falls back to 115200 + READY when it does not answer,
  and leaves the client's idea of the baud matching the port.

On hardware the forced desync now recovers in 1.1 s instead of hanging.
On a loaded host the FT232R link at 921600 dropped frames three times in
a row on the same batch, and identical retries failed identically. A
retry now moves the data at 115200 (restoring it first if the link is
still fast); slow, but it gets the batch.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant