Low priority. Nothing depends on this; it exists so three DRAFT markers do not quietly harden into asserted mappings.
What needs reviewing
evals/EXTERNAL_BENCHMARKS.md catalogues three published benchmarks (added in #118). Each is listed against an OWASP entry, and that column is a reading of the benchmark's stated scope, not a reviewed mapping:
| Benchmark |
Entry as drafted |
What it actually measures |
| LongPIBench |
LLM01 Prompt Injection |
Prompt injection in long-context settings — peer review, resume screening, code review, email summary; synthetic and real-world datasets |
| GenIaC-SecBench |
LLM10 Improper Output Handling |
Security of LLM-generated Infrastructure-as-Code against a human baseline; 100 scenarios, 12 model configurations, 1,196 artifacts |
| TIER |
LLM01 Prompt Injection |
Behavioural safety across threat implicitness; 4 risk domains × 4 threat levels, 6-label scale, 2 LLM judges |
The two I would look at hardest
- GenIaC-SecBench → LLM10. Insecure generated IaC is arguably closer to LLM05 (Data and Model Poisoning) or an agentic entry depending on how the code reaches production. LLM10 was chosen because the benchmark scores the output the model hands downstream; a reviewer may disagree.
- TIER → LLM01. TIER measures refusal behaviour under jailbreak-style pressure. LLM01 is the closest existing entry, but "safety behaviour" is not the same thing as prompt injection, and this may be a case where no current entry fits and the honest answer is none.
LongPIBench → LLM01 is the least contentious of the three.
Please do not
- Remove the DRAFT markers until this is signed off — they are what stops the column being read as an assertion.
- Treat these as mappings in the crosswalk data. They live in a catalogue file, not in
data/entries/, and nothing downstream consumes them.
Suggested timing
Review this in the same session as the ISO 27001 schema-v2 template, since both are judgment calls about how a source is represented rather than implementation work.
Low priority. Nothing depends on this; it exists so three DRAFT markers do not quietly harden into asserted mappings.
What needs reviewing
evals/EXTERNAL_BENCHMARKS.mdcatalogues three published benchmarks (added in #118). Each is listed against an OWASP entry, and that column is a reading of the benchmark's stated scope, not a reviewed mapping:The two I would look at hardest
LongPIBench → LLM01 is the least contentious of the three.
Please do not
data/entries/, and nothing downstream consumes them.Suggested timing
Review this in the same session as the ISO 27001 schema-v2 template, since both are judgment calls about how a source is represented rather than implementation work.