Skip to content
View fanyingcz's full-sized avatar

Block or report fanyingcz

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
fanyingcz/README.md

Chengye Yuan

M.Sc. in AI and Entrepreneurship — HKUST (2026–2028) B.Eng. in Computer Science and Technology — ECUST (2022–2026)

I build LLM and agent systems, and the evaluation harnesses that keep them honest. My work sits between applied AI engineering and measurement: retrieval-augmented pipelines, structured outputs and tool calling, and benchmarks that test whether a model actually understands a situation rather than pattern-matching it.


Selected projects

vlm-world-knowledge-benchmark A counterfactual image-understanding benchmark for vision-language models. It separates "the model recognises the picture" from "the model can reason about what would happen if one variable changed", using paired relevant / irrelevant knowledge conditions to isolate where accuracy gains actually come from.

ai-work-order-system Natural-language maintenance ticket triage: unstructured repair reports become categorised, geocoded, worker-assigned work orders. A pure-LLM pipeline was not reliable enough, so the LLM is constrained by a rule engine at every step. Dockerised FastAPI backend, Vue front end, public demo.


Contact — cyuanam@connect.ust.hk

Popular repositories Loading

  1. vlm-world-knowledge-benchmark vlm-world-knowledge-benchmark Public

    Counterfactual image-understanding benchmark for Vision-Language Models: 558 questions over 329 images, 4 prompting conditions, 3 production VLMs.

    Python

  2. ai-work-order-system ai-work-order-system Public

    LLM + rule engine pipeline for property-maintenance ticket triage. Cuts handling time from ~10 min to ~20 s and lifts classification accuracy from 92% to 96.6%.

    Python

  3. fanyingcz fanyingcz Public

  4. fanyingcz.github.io fanyingcz.github.io Public

    HTML