AI evaluation · scientific fact-checking · knowledge integrity
I test AI systems, verify scientific claims and reconstruct public evidence.
My work combines model-behavior evaluation, primary-source research, multilingual scientific content quality, structured-data verification and technical product operations. I build and maintain research-information tools when a recurring evidence problem needs a practical system.
Italy · EU work authorization · open to worldwide relocation · international contracting
Portfolio · AI Evaluation CV · Evaluation Record · LinkedIn · Email
Adversarial and edge-case testing across chat, multimodal inputs, agentic tool use and indirect prompt injection. I focus on observable behavior, test variation, reproducibility notes and conservative reporting of evidence boundaries.
Primary-literature review, source verification and evidence synthesis for scientific scripts, articles, visualizations and public communication.
Public-source investigation, archival recovery, consumer-genomics privacy research, corporate-source reconciliation, claim-to-source auditing, content-governance review and structured biomedical evidence synthesis.
Information architecture, provenance rules, metadata quality, entity reconciliation, functional testing, release verification and maintenance of public research tools.
Metrics are dated snapshots rather than permanent claims. Live ranks, platform counts, directory records and collaborative projects may change.
Self-directed testing of instruction handling, policy boundaries and edge cases across four interaction surfaces:
- chat and multi-turn behavior;
- multimodal and visual inputs;
- agentic tool use;
- indirect prompt injection.
I vary the interaction path, preserve reproduction notes and distinguish directly observed behavior from platform labels, independent verification and model-wide conclusions.
Evaluation methodology and public record · Dated evidence snapshot · AI Evaluation CV
An open-source directory of ways individuals can participate in scientific research by contributing data, biological samples, genomes, time or other forms of participation.
I founded the project and defined its:
- inclusion and exclusion criteria;
- source-verification workflow;
- provenance and update fields;
- metadata and entity model;
- licensing boundaries;
- human- and machine-readable interfaces;
- release-testing and maintenance process.
Live directory · Source repository · Project statistics
I work as an independent contractor with Entropy for Life, an established Italian science-communication brand.
My contributions include primary-literature research, scientific fact-checking, English-to-Italian localization, scripts, data analysis, visualizations, slides, short-form content, selected thumbnail work and website operations.
I also designed and built entropyforlife.it in WordPress and manage its publishing, responsive design and hosting environment.
Official work and attribution record
Notandia develops tools for identifying research-integrity signals across literature-search and reference-management workflows.
The currently released products identify MDPI references across browser surfaces and Zotero while avoiding ambiguous title-only matching.
My role covers product requirements, expected behavior, false-positive analysis, functional testing, release verification, deployment and ongoing maintenance. Implementation is AI-assisted.
Browser extension · Zotero plugin
A Telegram bot that converts non-English Wikipedia links into their corresponding English-language articles.
The service supports private chats, groups and inline use. Its operational architecture includes AWS Lambda deployment, durable processing, state management, monitoring and automated releases.
I specified the behavior and deployment requirements, tested user workflows, diagnosed production failures and maintain the service through AI-assisted implementation.
Use the bot · Source repository
My public work includes:
- consumer-genomics privacy and corporate-source reconciliation;
- contested-policy research;
- web-archive and missing-document recovery;
- legally sensitive chronology construction;
- scientific source-quality and bibliometric review;
- biomedical literature synthesis and taxonomy design;
- multilingual technical and scientific localization.
Each published case links to attributable public records and states what the evidence does—and does not—establish.
Separate observation from interpretation. Document what was directly observed before making broader claims.
Make assumptions visible. State the definitions, exclusions and judgments that affect the result.
Preserve provenance. Keep outputs connected to inspectable sources, records, revisions and dates.
Test behavior, not only structure. Verify edge cases, failure paths, releases and operational outcomes.
Plan for maintenance. Treat updates, corrections, deployment recovery and documentation as part of the work.
I am code-literate and work across:
- requirements and behavioral specifications;
- codebase and output inspection;
- Git and GitHub workflows;
- JSON, REST APIs and structured metadata;
- functional and regression testing;
- deployment diagnosis;
- WordPress and Cloudflare Pages operations;
- AWS Lambda and GitHub Actions workflows;
- provenance, taxonomy and validation rules.
Much of the software implementation in my projects is AI-assisted. I personally define the problem and acceptance criteria, inspect structure and behavior, test outputs and releases, diagnose deployment problems and maintain the resulting systems.
I do not present this work as evidence of independent software-engineering depth where that has not been demonstrated.
I am particularly interested in roles involving:
- AI evaluation and safeguards operations;
- model behavior and adversarial quality assurance;
- trust, safety and knowledge integrity;
- scientific AI and research-data quality;
- evaluation data, grading and human-feedback operations;
- scientific fact-checking and evidence verification;
- multilingual scientific content quality.
ORCID · European Nucleotide Archive · Tableau Public · Flourish
Mario Marcolongo Italy · Italian and EU citizen Italian native · English C1




