Skip to content
View RealJasonHu's full-sized avatar

Block or report RealJasonHu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
RealJasonHu/README.md
Jason Hu — AI research and open-source engineering

I build tools that turn hard-to-debug failures into reproducible, inspectable evidence.
把难以定位的问题,变成更小的复现案例、可执行测试与可审查的证据。

Selected work · Research · Upstream contributions · Engineering approach · All repositories

Focus: AI reliability Focus: world models Focus: multimodal systems Focus: developer tools

Selected work / 原创项目

Three open-source tools, connected by one idea: make a failure concrete enough to reproduce, inspect, and fix. These are early-stage tools with documented boundaries, not claims of production maturity or general model accuracy.

01 / ReproCut — smaller journeys, the same bug

Reduce a failing browser journey to a verified reproducer and a runnable regression test.

ReproCut runs deletion experiments in fresh Chromium contexts and keeps a shorter action sequence only when the same declared failure still reproduces. It exports a replayable journey, a Playwright regression test, an offline evidence report, and a repair brief for a developer or coding agent.

Actual ReproCut browser report showing a reduced journey and replay evidence

Recorded demo: 12 → 3 actions · 36 browser replays · 75% fewer steps. One deterministic fixture, not a cross-application benchmark. Verified reductions are single-deletion minimal under the observed replay conditions, not guaranteed globally shortest.

JavaScript Node.js Playwright Delta debugging

Source & quick start · Demo artifacts & releases · 中文介绍

Visual regression review, backed by inspectable evidence.

Actual RenderWitness report from captured browser fixtures, using the offline metric-only provider

Capture browser screenshots, isolate changed regions, and review the result with side-by-side, blend, and diff views.

Includes scenario suites, ignore regions, explicit CI gates, and HTML / JSON / Markdown / JUnit exports. Optional VLM analysis stays separate from deterministic pixel evidence.

The illustrated demo is metric-only; it does not establish real-VLM accuracy. Alpha tool.

Python Playwright VLM adapters CI

Source & demo · CI guide · 中文

Find, shrink, and replay world-model failures.

DreamFuzz — property-based testing for world models

Search for action sequences that expose model-vs-reference rollout drift, minimize the counterexample, and export a self-contained HTML replay.

Uses explicit behavioral properties, paired open-loop rollouts, reproducible seeds, and an adapter contract for custom targets. The core has no third-party runtime dependencies.

Alpha diagnostic tool, not a model-training framework or a general evaluation leaderboard.

Python Property-based testing World models

Source · Live replay · 中文

Research / 研究方向

I am interested in reliable multimodal systems, long-video understanding, and learning from limited supervision. My focus is on making model behavior testable, evidence traceable, and evaluation reproducible.

Public software is listed above. Research manuscripts, results, and code are linked only when cleared for public release.

Upstream contributions / 开源协作

I also work on focused bug fixes, regression tests, documentation, and code review in existing projects. Original projects, submitted patches, and review contributions are listed separately.

Merged

Project Contribution Evidence
Podman Desktop Fixed navigation resize-handle layering beneath the welcome overlay, with a Playwright hit-test regression. Merged PR #18998

Submitted pull requests

Project Contribution Evidence
Instructor Preserve GenAI text-part metadata during templating, including thought signatures, without mutating caller-owned content. PR #2608
Stable World Model Fix single-step iCEM planning so the existing horizon-one sampling path runs without an index error. PR #323
OpenPI Preserve histogram counts when normalization ranges expand, with regression coverage for quantile estimates. PR #1041
tqdm Fix notebook progress bars with CSS widths while preserving widget layout and numeric-width behavior. PR #1823
LiteLLM Preserve caller-provided spend metadata on authentication failures, with safe parsing and regression coverage. PR #38493
MuJoCo Clarify regularized-friction creep, the limits of impratio, and NoSlip tradeoffs in the documentation. PR #3526

Status checked on September 8, 2026: the six PRs above were open. The linked upstream threads are the source of truth for later changes.

Code review

LeRobot: reviewed the intermediate-prediction contract for world-model policies and reported regression risks around evaluation flag propagation and non-image outputs entering the video path. Review on PR #3757.

Engineering approach / 工程取向

Principle What it looks like in the work
A concrete failure before a broad claim A smaller browser journey, a model counterexample, or a numbered visual region—not just an aggregate score.
Evidence separate from interpretation Deterministic measurements and replay artifacts remain inspectable; model recommendations carry explicit limits.
Regression coverage with the fix Focused tests document the failure mode and guard the behavior being changed.
Useful handoffs Runnable tests, structured JSON, portable reports, and documentation that another developer can use.

Toolbox / 工具箱

Python JavaScript Node.js Playwright pytest GitHub Actions


Jason Hu
AI reliability · World models · Multimodal systems · Open source

Pinned Loading

  1. BerriAI/litellm BerriAI/litellm Public

    The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…

    Python 60.2k 12k

  2. huggingface/lerobot huggingface/lerobot Public

    🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning

    Python 28k 5.8k

  3. google-deepmind/mujoco google-deepmind/mujoco Public

    Multi-Joint dynamics with Contact. A general purpose physics simulator.

    C++ 15.5k 1.8k

  4. podman-desktop/podman-desktop podman-desktop/podman-desktop Public

    Podman Desktop is the best free and open source tool to work with Containers and Kubernetes for developers. Get an intuitive and user-friendly interface to effortlessly build, manage, and deploy co…

    TypeScript 8.1k 573