Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
49 changes: 49 additions & 0 deletions CHANGE_LOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,55 @@

This file records completed project work in chronological order.

## 2026-08-15

- Unblocked full-workbook conversion of the 2024 Canada FABLE Calculator workbook (Zenodo record
14755928, `2024_Open_FABLECalculator.xlsx`): the extract, graph, translate, infer-contract, and
generate steps now all complete cleanly (410,299 formulas translated, 0 untranslated, 0 output
closure blockers, `generated=True`) for both the pristine and edited workbook copies.
- Implemented static `INDIRECT(ADDRESS(ROW() +- k, COLUMN() +- k))` resolution (used by 55 cells as
a "value of the cell directly above" pattern): `ROW`/`COLUMN` parse to the formula cell's
coordinates, `ADDRESS` folds to a literal address string, `INDIRECT` resolves a static address to
a cell reference; `static_indirect_cell_reference` in `references.py` adds the resolved target to
`FormulaRecord.raw_references` so the dependency graph creates the execution edge.
- Added `repair_corrupted_structured_references` in `references.py` for the source-level defect
where a structured reference appears with a duplicated table prefix
(`name[] name[[#This Row],[Column]]`, 330 cells); the dangling prefix is dropped during
tokenization and raw-reference extraction while `FormulaRecord.raw_formula` keeps provenance.
- Added `ADDRESS` and `INDIRECT` to `SUPPORTED_FUNCTIONS`; non-static INDIRECT/ADDRESS forms report
`unsupported_function` instead of translating.
- Added regression tests: repaired corrupted structured reference, static INDIRECT translation,
non-static INDIRECT rejection, and static INDIRECT raw-reference extraction. Full test run:
212 passed, 1 skipped.
- Executed the generated 2024 Canada FABLE model end to end (`calculate({})`, 10,274 outputs,
~100 s) and eliminated all 349 runtime output errors found during scenario execution.
- Hardened generated runtime semantics in `generation.py` so invalid workbook arithmetic and math
return Excel-faithful error strings instead of raising Python exceptions:
`_sf_arith` now coerces numeric strings, treats blank as zero, returns `#DIV/0!` for division by
zero, and `#NUM!` for negative-base non-integer `^`; ordering comparisons (`>`, `>=`, `<`, `<=`)
render through a new coercion-safe `_sf_compare` (Excel number<text ordering, numeric-string
coercion); `VALUE`/`NUMBERVALUE` return `#VALUE!` and `LN` returns `#NUM!` instead of raising.
- Fixed `_sf_index` so `INDEX` resolves `_SfRangeView` arrays (single-column and multi-column row
reconstruction) instead of treating the whole range as one scalar row; out-of-range rows return
`#REF!`.
- Replaced raw `sum(_sf_flatten(...))` SUM rendering with a `_sf_sum` helper that propagates error
values and ignores text cells, and made `_sf_average` error/text safe (`#DIV/0!` on empty),
fixing `TypeError: unsupported operand type(s) for +: 'int' and 'str'` crashes from text/error
cells inside large aggregation ranges.
- Re-ran the scenario after regeneration: all 10,274 generated outputs now compute with zero errors;
a 300-output sample compared against workbook cached values matched 94% with small relative
deviations (0.05%-0.3%) consistent with Excel iterative-circular-calculation tolerance, not
systematic translation defects.
- Resolved the `static_circular_dependency` investigation (1,741 warnings): the cycles are phantom
whole-column SUMIFS ranges (e.g. `SUMIFS(calc_land_cor[CalcForest], calc_land_cor[Year], year-5)`)
whose year-offset criteria exclude the formula's own cell; lazy sum-range evaluation means they
never form runtime cycles, and the scenario run produced no circular-dependency errors. Decision:
keep these as warnings and document the phantom-cycle rationale rather than attempting automatic
cycle-breaking.
- Added regression tests for Excel-faithful arithmetic errors, numeric-string coercion in
arithmetic and comparisons, `LN`/`VALUE` error strings, `INDEX` over range views, and SUM error
propagation/text-skipping. Full test run: 217 passed, 1 skipped.

## 2026-07-02

- Activated Phase 38 on `feature/p38-matrix-generated-model-evidence`, created parent issue #243 and
Expand Down
4 changes: 2 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ dev = [
"formulas",
"pandas>=2",
"pytest>=8",
"ruff>=0.8",
"ruff>=0.8,<0.16",
"sphinx>=7",
"sphinx-rtd-theme>=2",
"twine>=5"
Expand All @@ -77,7 +77,7 @@ oracle = [
"formulas"
]
quality = [
"ruff>=0.8"
"ruff>=0.8,<0.16"
]
release = [
"build>=1.2",
Expand Down
32 changes: 27 additions & 5 deletions src/modelwright/extraction.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,9 @@
from openpyxl import load_workbook
from openpyxl.formula.tokenizer import Tokenizer
from openpyxl.utils.cell import get_column_letter, range_boundaries
from openpyxl.worksheet.formula import ArrayFormula

from modelwright.references import repair_corrupted_structured_references, static_indirect_cell_reference


JsonValue = str | int | float | bool | None | list[Any] | dict[str, Any]
Expand Down Expand Up @@ -474,12 +477,13 @@ def _extract_sheet_cells(

cell_ref = _cell_ref(worksheet.title, coordinate)
if cell.data_type == "f":
formula = _extract_formula(cell_ref, str(cell.value), cached_value)
formula_value = _formula_cell_value(raw_value)
formula = _extract_formula(cell_ref, formula_value, cached_value)
records.append(
CellRecord(
cell_ref=cell_ref,
kind="formula",
raw_value=_json_value(raw_value),
raw_value=_json_value(formula_value),
data_type=cell.data_type,
cached_value=_json_value(cached_value),
formula=formula,
Expand Down Expand Up @@ -524,8 +528,9 @@ def _progress(progress: Callable[[str], None] | None, message: str) -> None:
progress(message)

def _extract_formula(cell_ref: str, raw_formula: str, cached_value: JsonValue) -> FormulaRecord:
tokenized_formula = repair_corrupted_structured_references(raw_formula)
try:
tokenizer = Tokenizer(raw_formula)
tokenizer = Tokenizer(tokenized_formula)
except Exception as error:
return FormulaRecord(
raw_formula=raw_formula,
Expand All @@ -544,6 +549,9 @@ def _extract_formula(cell_ref: str, raw_formula: str, cached_value: JsonValue) -
raw_references = tuple(
token.value for token in tokenizer.items if token.type == "OPERAND" and token.subtype == "RANGE"
)
static_indirect_reference = static_indirect_cell_reference(cell_ref, raw_formula)
if static_indirect_reference is not None and static_indirect_reference not in raw_references:
raw_references = raw_references + (static_indirect_reference,)
functions = tuple(
token.value[:-1].upper() for token in tokenizer.items if token.type == "FUNC" and token.subtype == "OPEN"
)
Expand Down Expand Up @@ -636,6 +644,20 @@ def _json_value(value: Any) -> JsonValue:
return str(value)


def _formula_cell_value(value: Any) -> str:
"""Return the formula text for a formula cell value.

OpenPyXL stores array formula cells (and similar formula objects) as wrapper
objects that expose the formula text through their ``text`` attribute instead
of as plain strings.
"""
if isinstance(value, str):
return value
if isinstance(value, ArrayFormula):
return value.text if value.text else str(value)
return str(value)


def _is_external_reference(reference: str) -> bool:
return "[" in reference and "]" in reference and ("." in reference.split("]", 1)[0] or "!" in reference)

Expand All @@ -657,7 +679,7 @@ def _bracketed_parts(reference: str) -> tuple[str, ...]:
if character == "]":
depth -= 1
if depth == 0:
part = "".join(current)
part = "".join(current).strip()
current = []
if part.startswith("[") and part.endswith("]"):
parts.extend(_bracketed_parts(part))
Expand All @@ -672,4 +694,4 @@ def _bracketed_parts(reference: str) -> tuple[str, ...]:


def _clean_structured_selector(selector: str) -> str:
return selector.removeprefix("@").replace("''", "'")
return selector.strip().removeprefix("@").replace("''", "'")
117 changes: 111 additions & 6 deletions src/modelwright/formulas.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,9 @@
from modelwright.extraction import CellRecord
from modelwright.graph import DependencyEdge, DependencyGraph
from modelwright.references import WorkbookReference
from modelwright.references import cell_reference_coordinates
from modelwright.references import normalize_reference
from modelwright.references import repair_corrupted_structured_references


JsonValue = str | int | float | bool | None | list[Any] | dict[str, Any]
Expand All @@ -26,21 +28,31 @@
{
"AND",
"AVERAGE",
"AVERAGEIF",
"AVERAGEIFS",
"CONCATENATE",
"COUNTIF",
"COUNTIFS",
"IF",
"IFERROR",
"IFNA",
"INDEX",
"LN",
"MATCH",
"MAX",
"MIN",
"MINIFS",
"NUMBERVALUE",
"OR",
"OFFSET",
"ROUND",
"SUM",
"SUMIF",
"SUMIFS",
"VALUE",
"VLOOKUP",
"ADDRESS",
"INDIRECT",
}
)
SUPPORTED_OPERATORS = frozenset({"+", "-", "*", "/", "^", "&", ">", ">=", "<", "<=", "=", "<>", "(", ")", ","})
Expand Down Expand Up @@ -343,6 +355,8 @@ def _parse_primary(self) -> FormulaExpressionNode:
return FormulaExpressionNode.literal(token.value)
if token.kind == "logical":
return FormulaExpressionNode.literal(token.value == "TRUE")
if token.kind == "error":
return FormulaExpressionNode.literal(token.value)
if token.kind == "reference":
return FormulaExpressionNode.reference_to(self._resolved_reference(token.value))
if token.kind == "identifier":
Expand All @@ -357,6 +371,24 @@ def _parse_function_call(self, function_name: str) -> FormulaExpressionNode:
self._expect("(")
raw_function_name = function_name.upper()
function_name = _normalized_function_name(raw_function_name)
if function_name == "ROW":
if (token := self._peek()) is not None and token.value == ")":
self._advance()
return FormulaExpressionNode.literal(cell_reference_coordinates(self.cell.cell_ref)[0])
raise FormulaTranslationError(
"unsupported_function",
"ROW with arguments is not supported",
raw_function_name,
)
if function_name == "COLUMN":
if (token := self._peek()) is not None and token.value == ")":
self._advance()
return FormulaExpressionNode.literal(cell_reference_coordinates(self.cell.cell_ref)[1])
raise FormulaTranslationError(
"unsupported_function",
"COLUMN with arguments is not supported",
raw_function_name,
)
if function_name not in SUPPORTED_FUNCTIONS:
raise FormulaTranslationError(
"unsupported_function",
Expand All @@ -378,6 +410,10 @@ def _parse_function_call(self, function_name: str) -> FormulaExpressionNode:
self._expect(")")
if function_name == "OFFSET":
return _static_offset_reference(arguments)
if function_name == "ADDRESS":
return _static_address_reference(arguments)
if function_name == "INDIRECT":
return _static_indirect_reference(self, arguments)
return FormulaExpressionNode.function_call(function_name, tuple(arguments))

def _resolved_reference(self, raw_reference: str) -> WorkbookReference:
Expand Down Expand Up @@ -437,7 +473,8 @@ def _expect(self, value: str) -> None:

def _formula_tokens(raw_formula: str) -> tuple[_FormulaToken, ...]:
tokens: list[_FormulaToken] = []
for token in Tokenizer(raw_formula).items:
repaired_formula = repair_corrupted_structured_references(raw_formula)
for token in Tokenizer(repaired_formula).items:
if token.type == "WHITE-SPACE":
continue
if token.type == "FUNC" and token.subtype == "OPEN":
Expand All @@ -463,11 +500,8 @@ def _formula_tokens(raw_formula: str) -> tuple[_FormulaToken, ...]:
tokens.append(_FormulaToken("logical", token.value.upper()))
continue
if token.type == "OPERAND" and token.subtype == "ERROR":
raise FormulaTranslationError(
"unsupported_error_reference",
"formula contains an unsupported error reference",
token.value,
)
tokens.append(_FormulaToken("error", token.value))
continue
if token.type == "OPERAND" and token.subtype == "RANGE":
tokens.append(_FormulaToken("reference", token.value))
continue
Expand Down Expand Up @@ -530,6 +564,77 @@ def _literal_integer(node: FormulaExpressionNode) -> int | None:
return None


def _static_address_reference(arguments: list[FormulaExpressionNode]) -> FormulaExpressionNode:
if len(arguments) not in {2, 3, 4}:
raise FormulaTranslationError(
"unsupported_function",
"ADDRESS requires two to four arguments",
"ADDRESS",
)
row = _static_number(arguments[0])
column = _static_number(arguments[1])
if row < 1 or column < 1 or int(row) != row or int(column) != column:
raise FormulaTranslationError(
"unsupported_function",
"ADDRESS row and column must be static positive integers",
"ADDRESS",
)
address = f"{get_column_letter(int(column))}{int(row)}"
return FormulaExpressionNode.literal(address)


def _static_indirect_reference(parser: "_FormulaParser", arguments: list[FormulaExpressionNode]) -> FormulaExpressionNode:
if len(arguments) != 1:
raise FormulaTranslationError(
"unsupported_function",
"INDIRECT requires exactly one argument",
"INDIRECT",
)
address = _static_address_text(arguments[0])
if address is None:
raise FormulaTranslationError(
"unsupported_function",
"INDIRECT argument must be a static cell address",
"INDIRECT",
)
return FormulaExpressionNode.reference_to(parser._resolved_reference(address))


def _static_address_text(node: FormulaExpressionNode) -> str | None:
if node.kind == "literal" and isinstance(node.value, str):
return node.value
return None


def _static_number(node: FormulaExpressionNode) -> int | float:
if node.kind == "literal":
value = node.value
if isinstance(value, (int, float)) and not isinstance(value, bool):
return value
if node.kind == "unary":
(operand,) = node.operands
if node.operator == "-":
return -_static_number(operand)
if node.operator == "+":
return _static_number(operand)
if node.kind == "binary":
left = _static_number(node.operands[0])
right = _static_number(node.operands[1])
if node.operator == "+":
return left + right
if node.operator == "-":
return left - right
if node.operator == "*":
return left * right
if node.operator == "/":
return left / right
raise FormulaTranslationError(
"unsupported_function",
"function argument is not statically computable",
"",
)


def _shift_cell_reference(
reference: WorkbookReference,
*,
Expand Down
Loading
Loading