Public portfolio for adversarial AI evaluation, multimodal safety, model-behavior analysis, and evaluation quality systems.
Live site: jackiejay077.github.io
I work at the seam between red teaming and operations: design the evaluation, inspect the judgment, isolate the failure mechanism, and build the review layer that keeps recurring errors from becoming normal.
Selected operating evidence:
| Signal | Scope |
|---|---|
| Final-stage quality ownership | AI evaluation and red-team delivery workflows |
| 250+ QA audits per week | Monetization-integrity appeals and policy enforcement |
| 30% team error reduction | Structured audit findings and root-cause remediation |
| 800+ reviews per week | High-volume integrity and abuse queues |
| 40+ analysts trained | Onboarding, workflow documentation, and calibration |
| Multimodal red teaming | Text-to-image, image editing, and multi-reference evaluation |
- When Reassurance Overrides the Evidence A longitudinal safety case study examining reassurance override, context abandonment, and premature de-escalation.
-
Intent Classification Rubric An evidence-based method for classifying benign, ambiguous, adversarial, harmful, and indeterminate user intent.
-
Model Response Failure Taxonomy A reusable taxonomy for diagnosing failures across intent, context, grounding, refusal, and safety judgment.
-
Multimodal Evaluation Framework A structured method for locating failures across recognition, cross-modal integration, intent, grounding, and safety.
- A Refusal Is Not Automatically a Safe Response Why refusal is an output category while safety depends on reasoning quality, context, and proportional judgment.
- Pattern Completion as False Recall How inferred continuity can be presented as remembered context without verified retrieval.
The work in this repository emphasizes:
- conversation-level rather than prompt-level evaluation;
- evidence-based intent classification without collapsing ambiguity;
- separation of observable outcome, behavioral mechanism, and impact;
- combined interpretation of text, image, and multi-reference inputs;
- proportional safety judgment rather than refusal-counting;
- evaluator calibration, reproducibility, and root-cause remediation.
jackiejay077.github.io/
├── assets/
│ ├── css/
│ ├── icons/
│ ├── js/
│ ├── resume/
│ └── social/
├── case-files/
├── failure-nodes/
├── field-notes/
├── frameworks/
├── 404.html
├── index.html
├── robots.txt
└── sitemap.xml
The site is intentionally designed as a working evaluation environment rather than a generic portfolio template: dark operational UI, restrained teal status language, evidence-first document structure, and visible publication state.
All public examples are independently authored, synthetic, sanitized, paraphrased, or adapted from non-confidential work.
No proprietary datasets, internal policies, confidential prompts, restricted evaluation materials, or employer-owned taxonomies are reproduced here.
Jacqueline Jiang
AI Safety Analyst · Adversarial Evaluation · Multimodal Safety · Quality Operations