Skip to content
View SamCT86's full-sized avatar
💭
Open to Applied AI roles & contract builds
💭
Open to Applied AI roles & contract builds

Highlights

  • Pro

Block or report SamCT86

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SamCT86/README.md

Sarmad Tawfeek

AI-Native Product & Systems Builder | Stockholm, Sweden / Remote

I build and operate AI-native systems with an evidence-first approach: explicit acceptance criteria, real provider readback, fail-closed behavior, and verified outcomes instead of activity claims.

🏆 External performance proof — Yukon / Eigen Labs QSB
1.015B verified candidates/s on NVIDIA RTX 4090 · #47/1,515 scored submissions (top 3.1%) · 0.54% from the then-promoted best at the 2026-10-04 snapshot.
Inspect the upstream-verifiable benchmark evidence →

⚙️ External performance proof — Paradigm / ScoreBench Anthropic Take-Home
1,113 cycles · 132.7× cycle speedup vs the freshly verified starter · official provider rank #86 · ~top 15% of 565 participants at the 2026-10-05 snapshot.
Inspect the spoiler-safe official-score evidence →

Commercial product focus: MachineOutcome and Agent Cash Cow OS.

Start here if you're evaluating my work

  • MachineOutcome — inspect how I handle ambiguous external side effects before retrying. Relevant to agent/workflow systems that can create duplicate writes, deployments, payments or other irreversible effects.
  • Agent Cash Cow OS — start with the interactive Transaction Lab's buyer-first 15-second timeout walkthrough and guided failure run, inspect the public transaction-reliability source proof, then use the Forecast Evidence reference for the separate verification/runtime pattern behind the broader product direction.
  • Current engagement fit / next step — if one workflow is slow, expensive or unreliable, send 2–3 sentences by email. No technical brief or meeting is required to start, and no sensitive data should be sent yet. I return the next useful step; scope and price are agreed before anything is ordered. I do not claim paid adoption, customer ROI or market traction for MachineOutcome or Agent Cash Cow OS unless it is independently verified.

Other internal products and experiments are intentionally kept private and used as indie-hacking / R&D assets rather than presented as commercial portfolio products.

Current ways to start

These are the same bounded entry points presented on the live portfolio:

  • Workflow check — when you need to know whether a workflow problem is worth solving. You get a baseline, the biggest leak, estimated potential and a recommended next step.
  • Fix sprint — when the right problem is already clear. One bounded change is implemented against a pre-agreed metric and the outcome is verified.
  • Reliability review — when AI or integrations already run but you do not fully trust them. You get failure, duplicate-action and handoff testing plus a prioritized action list.

Start with 2–3 sentences about what is slow, expensive or unreliable. Portfolio · Email

LinkedIn · Email

Commercial products

MachineOutcome CI

MachineOutcome focuses on a critical automation problem: an attempted action is not the same thing as a verified outcome.

The public reference demonstrates reconciliation before retry, task/attempt-bound evidence, explicit VERIFIED | FAILED | UNKNOWN states, and refusal to blindly replay ambiguous external mutations.

Public proof: machineoutcome-case-study

Agent Cash Cow OS CI

Agent Cash Cow OS is the second commercial product track. Its public proof now covers two bounded surfaces: synthetic agent-commerce failure handling and forecast-evidence discipline for autonomous agents.

Interactive proof: Agent Cash Cow OS — Transaction Lab — a synthetic browser-only lab for timeout/readback, replay, authorization and outcome/settlement decisions. No account or real money is required.

The GitHub reference exposes a separate bounded engineering pattern for evidence, structured output, provider-state checks, cost/latency limits and fail-closed acceptance. The production OS, orchestration, payment/provider integrations, benchmark logic, live evidence and commercial controls remain private.

Public source proof: Agent Cash Cow OS — Transaction reliability proof — synthetic failure handling, proof receipts, deterministic tests and explicit public/private boundaries.

GitHub proof: Agent Cash Cow OS — Forecast Evidence reference

External open-source proof

Merged contributions to third-party repositories remain part of the engineering track record, but they are not commercial products.

  • gombit-dev/gombit — PR #541: made three test-only nil guards explicit so staticcheck could prove the following pointer dereferences unreachable; upstream CI passed, the maintainer reviewed the process contract, then approved and merged it.
  • gombit-dev/gombit — PR #534: fixed jobs retry --all reprocessing live re-failures, then addressed a maintainer-found peak-read regression with bounded pagination + regression coverage.
  • gombit-dev/gombit — PR #525: fixed unbounded crash redelivery so a poison job cannot keep rerunning past MaxAttempts.
  • gombit-dev/gombit — PR #521: fixed Dispatcher.DispatchAt mutating caller-owned option storage and added regression coverage.
  • LunaStev/binlayout — PR #39: connected the publishing result to the manual release flow.
  • LunaStev/binlayout — PR #35: added offline release-workflow regression coverage.

External benchmark proof

  • Yukon / Eigen Labs QSB pinning — independently verified RTX 4090 evaluation: 1,015,429,342 verified candidates/s. At the 2026-10-04 snapshot, the result ranked #47 of 1,515 scored submissions (top 3.1%), 0.54% below the then-promoted best. Yukon verified and scored the run successfully; it simply did not replace the promoted record. Inspect the evidence chain and exact upstream validation commit · Official benchmark
  • Paradigm / ScoreBench — Anthropic Take-Home: official 1,113-cycle result, 132.7× cycle speedup versus a freshly reproduced 147,734-cycle starter, provider rank #86 at the 2026-10-05 snapshot. The leaderboard contained 565 actual participant rows after excluding five reference baselines; ordered participant position was 85 / 565 (~top 15%). Inspect the spoiler-safe score, rank and verification evidence · Official challenge

How I work

  1. Define the real outcome and the evidence needed to prove it.
  2. Bound authority, side effects and failure states before execution.
  3. Use AI agents and software tools to implement the smallest causal solution.
  4. Test adversarially and read back the state that actually matters.
  5. Ship, reject or retry only from verified evidence.

AI gives me implementation leverage. I remain accountable for problem framing, system direction, orchestration, verification and final judgment.

Background

Before my current AI work, I spent years in investigation and evidence-heavy operational roles, with earlier studies in IT forensics and information security.

From 2015 to 2020, I also built and operated automated FX trading systems in live markets, where software behavior had direct financial consequences. A historical third-party performance record is available on Myfxbook.

Implementation environment

AI agents and coding tools · TypeScript / Node.js · Python / FastAPI · React / Next.js · Postgres / Supabase · OpenAI APIs · GitHub Actions

Pinned Loading

  1. agent-forecast-foundry-case-study agent-forecast-foundry-case-study Public

    A runnable AI agent reference that checks evidence, structured output, cost, latency and failure states before accepting a run.

    JavaScript

  2. machineoutcome-case-study machineoutcome-case-study Public

    A small reliability example that reads back real state before an agent retries an external action.

    JavaScript