Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .issueflows/03-solved-issues/issue164_original.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Issue #164: Allow for c.update()

- GitHub: https://github.com/jepegit/cellpy/issues/164
- Labels: enhancement, v2, cellpy2-stage4, cellpy2-stage5
- Epic: #783 (Epic L, live/incremental), Stage 2. Depends on: #780.

## Original description

After loading a file (e.g. a cellpyfile), it should be possible to update it by

```python
c = cellpy.get("cellpyfile.h5)
# the cellpy file contains the paths to the original raw files
c.update()
# only the new data will be loaded and processed
```

To achieve this, cellpy will need to have an easy way to find the last loaded
data and load from that. And update summaries from starting from the first new
step (or the last old if it is not complete) etc. This means that a smart way
of lookup and merging if several raw files are used is needed.
43 changes: 43 additions & 0 deletions .issueflows/03-solved-issues/issue164_plan.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
# Plan: #164 `CellpyCell.update()`

Confirmed approach (autonomous run under #783 stage 2; user asked to
"process the issues"). Builds on #778 (`update_core_data`), #779 (protocol),
#780 (`load_since` on four loaders).

## Approach

1. `CellpyCell.update(force=False, **loader_kwargs) -> bool` in
`cellpy/readers/cellreader.py`, before `merge`.
2. Change detection: fresh `FileID` size/mtime vs stored `raw_data_files`;
db sources always "changed".
3. Loader recovery from `data._provenance["source_type"]` via
`set_instrument` (cellpy-file loads carry the default tester).
4. Marker derived from `data.raw` (`_marker_from_raw`: both `row_count` and
`last_source_datapoint_num`, rewound to the last cycle start). No new
cellpy-file field.
5. Incremental path (single fid + native schema + harmonized raw + loader
matches `SupportsIncrementalLoad`): `load_since` → stamp `test_id`, align
dtypes → `_update_from_raw_rows` (`core.update_core_data` +
`_add_summary_extras` + `_refresh_scaled_summary_columns`).
6. Fallback full reload (`ValueError`/`LoaderError`, multi-file,
non-incremental loader): `from_raw` on all sources, restore
`meta_common` / `cycle_mode` / `cell_name`, `make_step_table`,
`make_summary(find_ir=...)`.
7. `_refresh_fid`: size/mtimes, `last_data_point`, `raw_data_files_length`.
8. `tests/incremental_support.incremental_update` delegates to the new engine.

## Tests (`tests/test_cell_update.py`, essential)

no-op, growth == full load, two growths (marker), FileID refresh, cellpy-file
round trip, single-cycle fallback keeps meta, force, no source raises,
non-incremental loader routes to full reload.

## Docs

`incremental-load-protocol.md` section for #164, `docs/agents/index.md`,
root `AGENTS.md` quick facts, HISTORY bullet, test-registry rows.

## Out of scope

Poll loop (#781), batch live refresh (#782), multi-file incremental merge
(falls back to full reload).
34 changes: 34 additions & 0 deletions .issueflows/03-solved-issues/issue164_status.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# Status: #164 `CellpyCell.update()`

- [x] Done

## Done

- `CellpyCell.update()` + helpers (`_raw_sources_changed`,
`_ensure_loader_for_update`, `_marker_from_raw`, `_update_incremental`,
`_update_from_raw_rows`, `_summary_has_ir`, `_refresh_fid`,
`_update_full_reload`) and module helpers `_frame_to_pandas`,
`_align_dtypes` in `cellpy/readers/cellreader.py`; `_load_marker` attribute
initialised in `__init__`.
- `tests/incremental_support.incremental_update` now delegates to
`CellpyCell._update_from_raw_rows` (the #778 oracle covers the shipped engine).
- `tests/test_cell_update.py`: 9 essential tests.
- Docs: `incremental-load-protocol.md` (#164 section), `docs/agents/index.md`,
root `AGENTS.md`, HISTORY `[Unreleased]`, test-registry.

## Verification

- `uv run pytest tests/test_cell_update.py tests/test_incremental_update.py tests/test_load_since.py` green.
- `uv run pytest -m essential` — see PR.

## Notes

- Branch `164-cell-update` stacked on `780-load-since` (PR #1101); PR base
is `780-load-since` until #1101 merges.
- Multi-file cells and non-incremental loaders take the full-reload path
(documented; the smart multi-file merge from the original issue is not
needed for the live use case, one file per running test).

## Remaining

- None for this issue. Stage 3: #781 (poll loop), #782 (batch live refresh).
49 changes: 49 additions & 0 deletions .issueflows/04-designs-and-guides/incremental-load-protocol.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,55 @@ expose `load_since`).
cycle-local rebase may differ from a full load. Only loader-made markers
carry the equality guarantee.

## `CellpyCell.update()` (#164)

The public consumer of the protocol, in `cellpy/readers/cellreader.py`
(section "incremental refresh"). Returns `bool` (frames changed).

- **Change detection** compares a fresh `FileID(fid.full_name)` size and
mtime with the stored `raw_data_files` entry (same stats
`check_file_ids` uses). Databases (`is_db`) and unreadable stats always
count as changed. `force=True` skips the check.
- **Loader recovery.** A cell loaded from a cellpy-file carries the config
default tester; `data._provenance["source_type"]` (persisted) names the
loader that read the raw. `update()` calls `set_instrument` from it (and
forwards `**loader_kwargs`, e.g. `model=`, since the model is not
persisted). Decision: no new cellpy-file field.
- **Marker without state.** The marker is not persisted either. When the
cell has no in-memory `_load_marker`, `_marker_from_raw` derives one from
`data.raw` with **both** seek fields filled (`row_count` = index of the
first row of the last cycle; `last_source_datapoint_num` = the datapoint
before it), so text and arbin loaders each find their field. Rewinding to
the cycle start mirrors the loaders' own policy above. Alternative
rejected: storing the marker in the cellpy-file (extra schema, and the
derived one is exact for loader-made markers anyway).
- **Incremental path** only when: one raw file, `native_schema`,
`config.reader.use_harmonized_raw`, and
`isinstance(loader_class, SupportsIncrementalLoad)`. The chunk gets
`test_id = active_test_id` and its dtypes cast to the existing raw's
(`_align_dtypes`; harmonize can yield Int32 where raw has Int64), then
goes through `_update_from_raw_rows` → `core.update_core_data` with the
by-value inputs cellpy owns (`nom_cap_abs`, current factor, raw limits),
`_add_summary_extras`, and `_refresh_scaled_summary_columns`.
`find_ir` follows whether the current summary has `ir_charge`.
- **Fallback = full reload** on `ValueError` / `LoaderError` from the
incremental path (typically core refusing a chunk whose start is at or
before the first kept row, i.e. a single-cycle head), for multi-file
cells, and for non-incremental loaders. `from_raw` on all recorded
sources, then `meta_common` (deep copy), `cycle_mode`, and `cell_name`
are restored before `make_step_table()` / `make_summary(find_ir=...)`.
- **FileID refresh** after either path: size / mtimes, `last_data_point`
(max datapoint), `raw_data_files_length[-1]`. A second `update()` on the
same file is then a no-op.
- `tests/incremental_support.incremental_update` (the #778 oracle) now
delegates to `CellpyCell._update_from_raw_rows`, so the equality tests
cover the shipped engine.

## Link

Design §3 in `cellpy-design-and-development/active/cellpy2-live-incremental-design.md`.
Tests: `tests/test_load_since.py` (#780), `tests/test_cell_update.py` (#164).
Consumers: `live.py` poll loop (#781), batch live refresh (#782).
## Link

Design §3 in `cellpy-design-and-development/active/cellpy2-live-incremental-design.md`.
Expand Down
9 changes: 9 additions & 0 deletions .issueflows/04-designs-and-guides/test-registry.md
Original file line number Diff line number Diff line change
Expand Up @@ -200,6 +200,15 @@ current issue**. `/iflow-doctor` may audit the whole suite against this table.
| tests/test_load_since.py::test_arbin_res_load_since_seeks_on_datapoint_and_rewinds_to_cycle_start | yes | yes | arbin_res.load_since | #780 | skip without mdbtools |
| tests/test_load_since.py::test_arbin_sql_load_since_filters_on_datapoint | yes | yes | arbin_sql.load_since | #780 | mocked `_query_sql` |
| tests/test_load_since.py::test_arbin_sql_query_gets_a_datapoint_clause | yes | yes | arbin_sql._query_sql | #780 | SQL clause on the fully qualified table |
| tests/test_cell_update.py::test_update_on_unchanged_source_is_a_noop | yes | yes | CellpyCell.update / _raw_sources_changed | #164 | size+mtime unchanged → False |
| tests/test_cell_update.py::test_update_after_growth_equals_full_load | yes | yes | CellpyCell.update (incremental) | #164 | head + grown tail == full cellpy.get |
| tests/test_cell_update.py::test_update_twice_tracks_the_marker | yes | yes | CellpyCell._marker_from_raw / _load_marker | #164 | two growths; marker rewinds |
| tests/test_cell_update.py::test_update_refreshes_file_id | yes | yes | CellpyCell._refresh_fid | #164 | size / last_data_point / lengths; second update no-op |
| tests/test_cell_update.py::test_update_after_cellpy_file_round_trip | yes | yes | CellpyCell._ensure_loader_for_update | #164 | loader from provenance; mass kept |
| tests/test_cell_update.py::test_update_falls_back_to_full_reload_and_keeps_meta | yes | yes | CellpyCell._update_full_reload | #164 | single-cycle head → core rejects → reload |
| tests/test_cell_update.py::test_update_force_reloads_an_unchanged_source | yes | yes | CellpyCell.update(force=True) | #164 | |
| tests/test_cell_update.py::test_update_without_raw_source_raises | yes | yes | CellpyCell.update | #164 | NoDataFound |
| tests/test_cell_update.py::test_update_uses_full_reload_for_non_incremental_loader | yes | yes | CellpyCell.update (protocol gate) | #164 | loader without load_since → full reload |
| tests/test_dbreader.py::test_missing_column_warns_once | yes | yes | readers.dbreader.Reader._pick_info | #1008 | warn-once per missing header |
| tests/test_dbreader.py::test_nom_cap_specifics_column_reaches_pages | yes | yes | batch._dbengine._create_pages_dict | #1008 | db value → pages |
| tests/test_dbreader.py::test_simple_db_engine_skip_file_search_excel_reader | yes | yes | batch._dbengine.simple_db_engine / find_files | #1017 | skip_file_search frames one row per cell |
Expand Down
3 changes: 3 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -374,6 +374,9 @@ Quick facts:
`kind=raw` lists raw and sets `needs_metadata`).
- Metadata peek (no frames): `cellpy.read_meta(path)` → dict with `cell` / `tests`.
- Ingestion form fields: `cellpy.instrument_meta_schema(instrument)` → `fields` / `units`.
- Live/running test: `c.update()` re-reads only the new raw rows (incremental
loaders) or reloads fully, returns `True` when frames changed; `False` if
the raw file did not change on disk. Also works after `cellpy.get(".cellpy")`.
- Frames: `c.data.raw` / `.steps` / `.summary`; columns via `c.schema.*`.
After a raw load, each cycle's raw capacity starts at 0. A forgotten tester
reset that 1.x plotted as doubled capacity is rebased on load for every
Expand Down
7 changes: 7 additions & 0 deletions HISTORY.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,13 @@
next marker, which rewinds to the start of the last cycle so cycle-local
capacity normalisation stays correct. Other loaders stay full-read. (#780)

* `CellpyCell.update()`: refresh a cell from a raw file that grew. Detects
changes from file size/mtime, reads only the new rows for incremental
loaders (arbin_res, arbin_sql, neware_txt, maccor_txt) and appends them
through cellpy-core, otherwise reloads fully while keeping mass, area,
nominal capacity, cycle mode and cell name. Works on cells loaded from a
cellpy-file. Returns `True` when the frames changed. (#164)

## [2.1.5.post6] - 2026-09-25

* `summary_collector(...).plot()` keeps a lone charge or discharge series
Expand Down
Loading
Loading