Skip to content
@aims-foundations

AI Measurement Science

Popular repositories Loading

  1. torch_measure torch_measure Public

    A package for AI measurement science

    Python 19 32

  2. aims aims Public

    Python 16 3

  3. reeval reeval Public

    Reliable and Efficient Model-based Generative Model Evaluation

    Jupyter Notebook 8 1

  4. fantastic-bugs fantastic-bugs Public

    Fantastic Bugs and Where to Find Them in AI Benchmarks

    Python 5 2

  5. safety-irt safety-irt Public

    A Measurement Analysis of Multilingual Safety Evaluation (Accepted to COLM 2026)

    Python 2 1

  6. dynamic-irt dynamic-irt Public

    Dynamic Measurement Models

    Python 1 1

Repositories

Showing 10 of 14 repositories
  • measurement-db Public

    Measurement Data Bank

    aims-foundations/measurement-db's past year of commit activity
    Python 0 CC-BY-SA-4.0 0 0 0 Updated Sep 18, 2026
  • aims-foundations/paiec_baseline's past year of commit activity
    Python 1 0 0 0 Updated Sep 16, 2026
  • evarium Public

    Generative Simulation of the AI Evaluation Ecosystem

    aims-foundations/evarium's past year of commit activity
    Python 0 0 0 0 Updated Sep 10, 2026
  • aims Public
    aims-foundations/aims's past year of commit activity
    Python 16 3 0 1 Updated Sep 7, 2026
  • dynamic-irt Public

    Dynamic Measurement Models

    aims-foundations/dynamic-irt's past year of commit activity
    Python 1 1 0 1 Updated Aug 17, 2026
  • benchmark-caliper Public

    Uplifting Human Decision Making in AI Evaluation by Automating Benchmark Validity Analysis

    aims-foundations/benchmark-caliper's past year of commit activity
    Python 1 2 0 1 Updated Aug 10, 2026
  • safety-irt Public

    A Measurement Analysis of Multilingual Safety Evaluation (Accepted to COLM 2026)

    aims-foundations/safety-irt's past year of commit activity
    Python 2 1 0 0 Updated Jul 20, 2026
  • torch_measure Public

    A package for AI measurement science

    aims-foundations/torch_measure's past year of commit activity
    Python 19 MIT 32 2 6 Updated Jul 4, 2026
  • redteam-measurement Public

    Long-form response matrices for adaptive AI red-teaming benchmarks (JailbreakBench, HarmBench, StrongREJECT, Do-Not-Answer), formatted to the aims-foundations/measurement-db schema.

    aims-foundations/redteam-measurement's past year of commit activity
    Python 0 0 0 0 Updated Jul 1, 2026
  • irsl Public

    Item Response Scaling Law

    aims-foundations/irsl's past year of commit activity
    Python 1 1 0 0 Updated May 26, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.