Skip to content

Explore interoperability with EveryEvalEver #10

Description

@yzeng58

Source\n\nLeshem Choshen (@LChoshen), an author of Every Eval Ever, commented on Yuchen's BenchPress post asking whether BenchPress data could be combined into EveryEvalEver.\n\n## Paper/project read\n\nEvery Eval Ever is a shared schema, converter/validation pipeline, and community Hugging Face datastore for evaluation result records. It standardizes source metadata, model information, generation/evaluation configuration, metric semantics, aggregate scores, and optional instance-level outputs.\n\nThis appears complementary to BenchPress: BenchPress maintains a curated model-by-benchmark score matrix and prediction layer, while EEE provides a broader interoperable record format and crowdsourced datastore.\n\n## Monthly maintenance action\n\nDuring a future monthly BenchPress maintenance cycle, evaluate whether to:\n\n- export BenchPress public scores to the EEE schema;\n- map BenchPress metadata fields to EEE source/model/config/result blocks;\n- document the relationship between BenchPress and EEE;\n- optionally open an upstream contribution/PR to the EEE datastore.\n

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions