Skip to content

GK7205V500 (FMC100 SPI-NAND): silent uncorrectable-ECC rots the UBIFS rootfs — same defect as #2285 #2519

Description

@widgetii

The Goke GK7205V500 family (incl. GK7205V510, e.g. Zenointel SD-2N-4G) exhibits the same silent-ECC SPI-NAND rootfs corruption first characterised on hi3516ev300 (hifmc100) in #2285. gk7205v500 uses the twin xmedia_fmc100 controller/driver, and the defect transfers 1:1. On nand-ultimate (rw UBIFS) the rootfs silently rots; on a sufficiently-degraded unit a fresh flash is corrupt on first mount.

Affected code (gk7205v500)

  • drivers/mtd/nand/xmedia_fmc100/fmc100.c:986 — chip->ecc.mode = NAND_ECC_NONE → nand_read_page_raw() (always reports 0 bitflips).
  • fmc100.c:637-651 — the only code touching mtd->ecc_stats is behind #ifdef CONFIG_XMEDIA_NAND_ECC_STATUS_REPORT.
  • br-ext-chip-goke/board/gk7205v500/gk7205v500.generic.config:950 — # CONFIG_XMEDIA_NAND_ECC_STATUS_REPORT is not set.
  • fmc100.h:189 — FMC100_ECC_ERR_NUM0_BUF0 = 0xc0; fmc100.h:191 — GET_ECC_ERR_NUM(i,reg)=(reg>>(i*8))&0xff.
  • DTS flash-memory-controller@10000000, reg = <0x10000000 0x1000>, <0x14000000 0x10000> — identical base to hi3516ev300.

Mechanism

  1. Silent: an uncorrectable ECC read returns garbage with mtd_read()=success; UBI never sees -EBADMSG, never scrubs; the driver never sets mtd->ecc_strength/ecc_step_size (sysfs ecc_strength=0). ecc_failures/corrected_bits=0 mean not implemented, not no errors.
  2. Engine: every UBIFS ubifs_write_master() (on each sync/commit) rewrites LEB 1/2 and those pages come back with unusable ECC parity — one new uncorrectable page per commit. (Why a UBIFS master write damages its block while an identical nandwrite does not was never explained in hi3516ev300 + Winbond W25N01GV: silent data-pattern UBIFS rootfs corruption (kernel corrupts its own writes; NAND HW proven good) — hifmc100 ECC bug #2285.)

Confirm on a live board

cat /sys/class/mtd/mtd3/ecc_strength         # 0  == bug present (should be 8)
devmem 0x10000000 32                         # FMC_CFG -> 0x00011823 (ECC_TYPE=8BIT, ON)
dd if=/dev/mtd3 bs=2048 skip=<PAGE> count=1 >/dev/null 2>&1
devmem 0x100000c0 32                         # ECC err: 0x0000FF00 = UNCORRECTABLE step 1  <-- signature

ECC layout: 2K page / 64B OOB, 8-bit BCH, step = 1040 data + 14 parity (stride 1054); failure is always step 1.

Deterministic reproducer (seconds)

Locate UBIFS LEB 1/2 (master nodes), snapshot their pages via devmem 0x100000c0, then:

for i in $(seq 0 19); do echo x > /root/c.$i; sync; done

→ one NEW 0x0000FF00 (uncorrectable) page appears per commit, in lockstep, on both master-node LEBs. Ruled out in #2285: worn silicon, data content, page position, program-disturb, the general UBIFS write path, and the ECC-report patch (reproduces on stock) — the signature is bad parity (0xFF, never 1–7 corrected).

Observed live (GK7205V510 / SD-2N-4G)

A freshly nand write-flashed gk7205v500-nand-ultimate rootfs.ubi (mtd3 0x600000, UBI# verified at flash) was already corrupt on first mount:

UBIFS error (ubi0:0): ubifs_check_node: bad CRC: calculated 0xa3ecf000, read 0xa626e930
UBIFS error (ubi0:0): bad node at LEB 101:57080        (directory node)

An earlier run: ubifs_iget: failed to read inode …, -117 → dead directory entry 'lib' → No working init found → kernel panic. nand bad showed no bad blocks in the root partition — consistent with silent, unmarked degradation.

Why the obvious patch is not a fix

Enabling ECC reporting (-EBADMSG on uncorrectable; the hi3516ev300 hifmc100-ecc-report.patch, a 1:1 HISI→XMEDIA rename for Goke) makes the fault visible so UBI could scrub a healthy board — but on a board that already has uncorrectable pages in the master/boot path it turns silent corruption into a hard mount-time panic (Cannot open root device ubi0:rootfs: error -74). Telemetry/prevention, not a cure; unproven/unmerged.

Harm-reduction (no cure today)

Read-only rootfs (squashfs / nand-lite), minimise sync/rw churn, keep rw state off NAND. Note ubinfo -a available PEBs: 0 (rootfs + rootfs_data fill the device) leaves UBI no spares to relocate.

Related: #2285 (hi3516ev300 / hifmc100 — same defect, full root-cause investigation).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions