Webcam, video file or iPhone → ARKit blendshape scores → MetaHuman face board → RigLogic.
Drives a MetaHuman face rig in Blender from an ordinary camera, and contains a differentiable reimplementation of RigLogic in PyTorch that makes the inverse problem — which control combination produced these landmarks? — tractable.
Status: working, actively developed, rough edges documented honestly below. The solver is experimental and its limits are measured, not hidden.
Epic's RigLogic is a compiled CPU library. It is fast to evaluate forward, but it is a black box: no gradients, no GPU, no batching. A numeric Jacobian needs N+1 sequential RigLogic calls per iteration — at 174 parameters that was 341 ms per iteration, which makes landmark-driven solving impractical.
Rewriting the evaluation chain as differentiable tensor ops removes the N+1 factor entirely. Autograd returns the full gradient in one backward pass:
| numeric Jacobian | differentiable rig | |
|---|---|---|
| cost per iteration | O(N) RigLogic calls | O(1) |
| measured gradient time | 341 ms | 4.5 ms |
| accuracy vs real RigLogic | — | 3.3e-07 max error |
That is not a 2x speedup, it is a change of complexity class. Accuracy is verified
against the real library in tests/test_riglogic_torch.py.
The reconstructed chain (recovered by probing the API, see
scripts/05_extract_behavior.py):
GUI (174)
| piecewise linear: raw = slope * gui + cut, gui in [from, to]
raw (263)
| PSD: weighted PRODUCTS of raw controls, appended after raw
control (808 = 263 + 545)
| 122 joint groups, each a dense (out x in) matrix
joint deltas (7830 = 870 x 9)
| direct gather
blendshape weights (782)
MediaPipe cannot be installed into Blender's Python; the protobuf/numpy conflict breaks Blender. So the capture side and the rig side are separate processes that talk over UDP:
[detector/.venv Python 3.11] [Blender 5.0 Python 3.11]
webcam / video meta_human_dna 0.5.4
| ^
mediapipe FaceLandmarker |
| 51 ARKit scores + head matrix |
| blender_addon/charface
+-- offline: takes/*.jsonl -------> charface.bake_take
+-- live: UDP :11111 --------> charface.live
|
core/ (mapping, filters, take schema — stdlib, shared)
core/ is imported by both sides, so the mapping logic lives in exactly one place.
ARKit pose --(Epic posemap)--> raw CTRL_expressions.* --(inverse GUI->raw)--> GUI axis
|
face board pose_bone.location ----------+
|
RigLogic --> bones + shape keys + masks
Raw controls are never written directly — the addon calls
mapGUIToRawControls() on every evaluation and regenerates raw values from GUI,
so anything you write to a raw control is overwritten in the same frame. Details
with file:line references: docs/api-notes.md.
Python 3.11 is required: MediaPipe publishes no wheels for 3.13/3.14, and the addon's DNA bindings are compiled for py311 too.
py -3.11 -m venv detector/.venv
detector/.venv/Scripts/python -m pip install -r detector/requirements.txt
curl -L -o detector/models/face_landmarker.task https://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/1/face_landmarker.taskpy -m venv .venv
.venv/Scripts/python -m pip install -r requirements-dev.txt
.venv/Scripts/python -m pip install --index-url https://download.pytorch.org/whl/cu128 torchOn Blackwell (RTX 50, sm_120) the cu128 wheel is mandatory. Torch from the
default index fails on these cards with no kernel image is available for execution on the device. Verify with a real kernel, not just is_available():
.venv/Scripts/python -c "import torch; x=torch.randn(512,512,device='cuda'); print((x@x).sum())"Torch is installed here deliberately and not borrowed from another environment:
core/riglogic_torch.py is a real dependency of facecap.
Install blender_addon/charface as an addon (or symlink it into scripts/addons).
It sits on top of Character DNA Pro / meta_human_dna, it does not replace them.
python scripts/install_addon.py --version 5.0Add --uninstall to remove it.
The mapping table and the solver's forward model are extracted from a MetaHuman DNA file and from Epic's ARKit posemap asset. That content is licensed by Epic and is not redistributed here. Generate it from your own character:
blender --background --python scripts/01_dump_dna_controls.py -- <head.dna> mapping/_generated/dna_controls.json
python scripts/02_build_mapping.py --posemap <posemap.json> --controls mapping/_generated/dna_controls.jsonFor the solver, additionally:
blender --background --python scripts/05_extract_behavior.py -- <head.dna> mapping/_generated/behavior.npz
blender --background --python scripts/04_extract_face_model.py -- <head.dna> mapping/_generated/face_model_landmarks.npz --vertices mapping/_generated/correspondence/correspondence.json
blender --background --python scripts/06_dump_reference.py -- <head.dna> mapping/_generated/riglogic_reference.npz 32Coverage: 51/51 MediaPipe channels plus tongueOut for iPhone. mouthClose
could not be derived from Epic's asset and is hand-written — source and confidence
bounds in docs/api-notes.md §8 and §11.
Run facecap.bat. Camera selection, sending to Blender, recording and live
channel bars in one window.
+-------------------------+----------------------+
| | Camera [0] [Scan] |
| preview | UDP 127.0.0.1:11111 |
| | [x] Preview |
| | [x] Filter |
| | [ ] Raw landmarks |
| | [ ] Write to file |
| | [ START ] |
| | face YES | 28 fps |
+-------------------------+----------------------+
| Mouth / jaw |
| mouthSmileRight ########## 0.96 (0.96) |
| jawOpen ### 0.27 (0.28) |
| Eye |
| eyeBlinkRight ###### 0.64 (0.77) |
| Brow / nose |
| browDownRight ### 0.37 (0.76) |
+------------------------------------------------+
The channel bars are the heart of calibration. Without seeing which channel fires how much when you make an expression, finding the dead one is guesswork.
The bars show deviation from neutral, not the raw value, and they are grouped by region. A single "top 8" list was useless, because MediaPipe reports high eye and brow values even on a neutral face. Measured on a real capture, subject sitting expressionless:
eyeSquintLeft 0.473 browDownRight 0.395 browDownLeft 0.305
eyeLookUpLeft 0.301 eyeBlinkLeft 0.290 eyeSquintRight 0.267
Eight channels permanently above 0.2, while mouth channels — genuinely 0.00 at
neutral — never entered the list at all. Eyes mislead too: eyeBlinkLeft starts at
0.290, so a raw 0.87 is really a 0.58 blink. The raw value in parentheses is kept
so you can see whether a channel is saturating.
Neutral is learned from the first 30 frames after start — hold a neutral face
then. If lighting or distance changes, hit Re-take Neutral.
Preview is capped at 15 fps (drawing every frame into Tk costs more than capture).
cv2.imshow is deliberately not used: HighGUI runs its own event loop and can
deadlock next to the Tk mainloop on Windows.
Produce a take from a video file:
detector/.venv/Scripts/python detector/detect.py video --input take.mp4 --out takes/take.jsonlLive:
detector/.venv/Scripts/python detector/detect.py live --udp 127.0.0.1:11111 --previewThen Live Face Capture in Blender (ESC to exit).
Two routes. Live recording is the normal one; baking is for re-trying a take.
facecap.bat→ START- Blender: Live Face Capture
- Put the playhead on the frame you want
- Press
Record to Timeline(turns red), play, press again
Keyframes are written while you play and the playhead advances. The panel shows the live count.
Timing comes from the wall clock, not the packet counter: packets arrive at 30 fps while the scene is usually 24 fps, so one keyframe per packet would slow the animation by 25%. When several packets land on the same scene frame, the axis value with the largest absolute value wins — taking the last one loses an 89 ms blink peak (verified headless: 30 packets → 25 keyframes, peak 0.9 preserved).
The scene range is only ever extended, never narrowed, so it cannot clip your existing animation. The real range is printed in the panel.
Useful for re-running the same take with a different calibration.
- In the GUI tick Write to file, START, perform, STOP
- Blender: N-panel → charface → Bake Take to Face Board → pick the .jsonl
| Option | Meaning |
|---|---|
Start Frame |
scene frame the animation starts on |
Neutral Calibration |
treats the first 30 frames as neutral and subtracts it. Leave on |
Single Channel |
only this ARKit channel (validate with jawOpen first) |
One-Euro Filter |
enable here if you did not filter while recording |
Mirror Left/Right |
the user's left is the character's right in a camera image |
The take is fitted to the scene fps by timestamp, not 1:1 frame to frame.
Measured: a webcam delivers 30 fps, Blender's default scene is 24 fps — 1:1 writing
turns a 37 s take into 46 s. Frame intervals are not constant either (12–50 ms
measured). Logic in core/take.py:plan_scene_frames, tests in tests/test_take.py.
On your first attempt put jawOpen in Single Channel. If the jaw opens, the
chain is sound — clear the field and bake again. Wiring everything at once hides
which layer is broken.
N-panel → charface → Head Calibration. Settings live on the scene, so they can be changed while the live stream runs and take effect immediately — as an operator property, every experiment would have meant stop/start.
- Live readout: yaw / pitch / roll in degrees, plus packet counter and face YES/NO. Without seeing how far each angle travels, finding the wrong axis is guesswork.
- Re-take Neutral: neutral is built from the first 30 frames; if you were not sitting straight then, the head stays permanently skewed. Look straight, press.
- Per-angle mapping: target bone axis +
Invert+Gainfor each angle. Gain 0 disables the angle, 2 doubles it. - Axis test: X / Y / Z keys rotate the head on that axis so you can see what
each one does.
Resetreturns to neutral.
Order that works: axis test first, then set the angle mapping, then fix direction
with Invert, and only then tune amount with Gain.
Channels do not reach 1.0. Measured on a real iPhone recording: eyeBlink
saturates at 0.917, mouthSmile at 0.865. Since the rig is built to open fully
at 1.0, the eye never fully closes without this step.
In the Face Calibration panel:
- Press
Learning Range - Perform every expression to its limit (blink, open jaw, smile, raise brows…)
- Turn it off
Save Ranges— otherwise it is lost when Blender closes
The profile is saved to //charface_profile.json (next to the scene, or in the
project root if the scene is unsaved) and loaded automatically when live capture
starts. Saving does not clobber the existing profile: gains and dead zones you set
by hand are preserved, only input_max is updated.
The operator only listens on UDP; it does not know who is sending.
| Source | Status |
|---|---|
scripts/replay_take.py |
✅ works — replays a recorded take in real time, fake detector |
detector/detect.py live |
✅ webcam — verified on real photographs |
| iPhone Live Link Face | ✅ supported — parser verified against 546 real packets |
The operator auto-detects the packet: JSON means webcam/replay, binary means Live Link Face. No setting, same button, same port.
| webcam (MediaPipe) | iPhone (ARKit) | |
|---|---|---|
| channels | 51 | 52 |
tongueOut |
no | yes |
mouthClose |
no (hand-written) | real measurement |
| stability | RGB estimate | TrueDepth sensor |
| eye gaze | from blendshapes | 6 separate channels |
Setup:
- App Store → Live Link Face (Epic Games, free). iPhone X or newer.
- Top-left gear → Live Link → Add Target → your PC's IP, port 11111
- Phone and PC on the same network
- Verify:
iphone_test.bat(orpython scripts/llf_probe.py) - Live Face Capture in Blender
The packet layout is not officially documented by Epic and community sources
contradict each other. The parser was run against 546 packets from a real
iPhone recording: all parsed, the fallback path was never taken. Details in
docs/api-notes.md §12. llf_probe.py still reports which path
was used — if it says scan, the layout differs on your iOS version.
N-panel → charface → Scene Performance. Unrelated to live capture, always usable.
The hair emitter is bound to a mesh that RigLogic deforms continuously, so every
deformation recomputes every strand. The real cost is not visible in the count
field — it is multiplied by child particles:
2000 strands x 25 children = 50,000 strands, each with 2^3 segments
The panel computes and shows this product. Lighten Hair lowers the viewport
values (display_percentage 10, child_percent 0, display_step 1) — measured:
50,000 strands → 200. Rendering is unaffected; render_step and
rendered_child_count are separate fields and are left alone. The backup is stored
in the particle settings datablock, so Undo restores the exact previous values
even across a save/reload.
With the addon's evaluate_dependency_graph flag on, RigLogic runs on every
depsgraph update (meta_human_dna/rig_logic.py:37) — selecting an object,
changing frame, editing particles all re-evaluate an 800+ bone rig. The addon only
disables this programmatically during bake/import and exposes it in no UI.
The panel lets you turn it off; the face then stops updating on its own and you
trigger it with Evaluate Now. Worth keeping off while doing hair or modeling work.
Preview Hz throttles the viewport update. It does not affect recording — keyframes come from the incoming values, not from what is on screen, so the animation records at full rate even at 15 Hz viewport. The panel shows measured numbers: rig ms, apply ms, and how many frames per second that allows.
The ARKit mapping works channel by channel: "I saw jawOpen 0.42, write 0.42 to the jaw control". The solver does the inverse: which control combination produces this landmark configuration? This is what MetaHuman Animator does.
With the default head.dna, the fit scale to Blender world came out as 0.010565
(should be 0.01 — a 5.7% error) with 2.75 mm residual. The same measurement with the
character's own DNA gives 0.009994 and 0.47 mm — 5.8x better. Modeling a
different head is the failure that happens silently and ruins everything downstream.
detector/.venv/Scripts/python detector/detect.py live --out takes/take.jsonl --landmarks --preview
python scripts/10_solve_take.py --take takes/take.jsonl--landmarks also writes all 478 raw landmarks (about 30x larger frames, hence not
the default). The first 30 frames of the take must be expressionless; neutral is
learned there and identity difference is cancelled against it.
The first correspondence was built with the head turned ~17°. MediaPipe found more landmarks (473 → 478), but landmarks on the far cheek were bound by a front-cast ray to vertices on the near side. Measured cost:
| turned head | neutral + tangent filter | |
|---|---|---|
| landmarks | 478 | 447 |
| DNA → Blender fit residual | 0.471 mm | 0.028 mm |
| left / right landmark balance | 176 / 260 | 202 / 207 |
Render Head now zeroes head rotation and the whole facial expression for the
duration of the render and restores them afterwards, whatever pose the character is
in. Landmarks whose ray hits the surface at more than ~70° are also culled
(facing < 0.35).
Landmark count is not the goal. Few and correctly bound beats many and wrong.
Error in mm, lower is better:
| frame | expression | baseline (no controls) | ARKit mapping | solver |
|---|---|---|---|---|
| 143 | smile | 5.360 | 7.139 | 3.335 |
| 187 | jaw open | 1.488 | 3.893 | 0.670 |
| 246 | blink | 1.291 | 1.887 | 1.100 |
| 775 | lip pucker | 1.847 | 2.917 | 0.838 |
The solver beats the ARKit mapping by a wide margin on every frame. ARKit comes out worse than the baseline because it finds the mouth correctly but injects fake motion across the rest of the face: 57 landmarks that move 0.7 mm in the observation are driven 4.9 mm, at cosine −0.19.
But the solver still explains only 32–51% of the motion. This is not an optimization problem: with regularization fully off, 2000 iterations and 136 controls, the result gets worse (3.438 vs 3.365 mm). The ceiling is in the model.
The open problem: transferring a real human's landmark motion onto a particular MetaHuman rig needs more than rigid alignment plus delta transfer. MetaHuman Animator solves it by first extracting an actor-specific DNA (identity solve), solving the expression on that rig, then retargeting. That step does not exist here.
Tried and did not work (recorded so it is not retried): solving pose jointly with
expression (solve(joint_pose=True)) — 3.475 → 4.026 mm on the smile frame; the pose
and expression gradients eat each other.
core/ |
Shared logic: mapping, one-euro filter, calibration, take schema. Stdlib-only except facemodel.py (numpy) and riglogic_torch.py (torch). |
detector/ |
MediaPipe side. Separate venv. |
blender_addon/charface/ |
Bake and live operators + panels. |
mapping/arkit_to_mh.json |
Generated mapping table. Not hand-edited. Not redistributed — see Install §4. |
mapping/_generated/ |
DNA dump, build report, introspection output. Not redistributed. |
gui.py |
Tk interface. Runs under detector/.venv; launched by facecap.bat. |
scripts/ |
Introspection, DNA dump, table generation, fixtures, UDP replay/monitor, addon install. |
tests/ |
Unit tests that do not require Blender. |
.venv/Scripts/python -m pytest tests/| Phase | Status |
|---|---|
| 0 — Discovery | ✅ docs/api-notes.md, with file:line references |
| 1 — Detector | ✅ verified — 51 scores + head quaternion from a real photograph |
| 2 — Mapping | ✅ 51/51 channels, 88 axes; mouthClose hand-written (§11) |
| 3 — Blender applier | ✅ live path verified in a real MetaHuman scene; bake operator not yet field-tested |
| 4 — Calibration / filter | ✅ head + face calibration tunable live, group gains, three-band one-euro |
| 5 — Export | uses the addon's own bake/export, no extra code |
| 6 — Derived controls | ✅ eyelid/gaze coupling, lid pressure, micro-saccade, channel range (§15) |
| 7 — Differentiable rig | ✅ RigLogic rewritten in PyTorch, 3.3e-07 error, 4.5 ms gradient (§14) |
| 8 — Landmark correspondence | ✅ 477/478 landmarks, barycentric, from the character's own head (§15) |
| 9 — Solver | ✅ two-stage (L1 + debias), 0.002 control error on synthetic tests (§16) |
| 10 — Identity separation | ✅ core/observation.py — observation/model neutral difference cancelled |
| 11 — Real-face trial |
This project interoperates with, but does not redistribute, the following:
- MetaHuman, RigLogic, the ARKit posemap asset (
PA_MetaHuman_ARKit_Mapping) and DNA files are © Epic Games and licensed under Epic's terms. No MetaHuman content is included in this repository;scripts/01–06regenerate what is needed from your own licensed assets. - Character DNA Pro /
meta_human_dnais a third-party Blender addon that must be installed separately. - MediaPipe (
face_landmarker.task) and Live Link Face are downloaded or installed separately under their own licenses.
The Turkish original of this document is kept at docs/README.tr.md.