# Pixel3dmm 3dface Recon Eval

> pixel3dmm-3dface-recon-eval

- Skill: `qhjqhj00/pixel3dmm-3dface-recon-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/pixel3dmm-3dface-recon-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/pixel3dmm-3dface-recon-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/pixel3dmm-3dface-recon-eval

---


# pixel3dmm-3dface-recon-eval

> Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction — Giebenhain et al. (2025) (arXiv:2505.00615, 2025)

## What this evaluates

This benchmark evaluates single-image 3D face reconstruction pipelines on two tasks: posed reconstruction (measuring geometric fidelity under diverse expressions) and neutral reconstruction (testing the ability to disentangle identity shape from expression). It probes a model's capacity to recover accurate facial geometry and surface normals from a single input image.

## Datasets

- **Pixel3DMM Benchmark** — total 441; splits: benchmark (441)

## Metrics

- `Chamfer Distance (L1/L2)` **(primary)** — range: other
  - Uni-directional distance from ground truth points to the nearest surface point on the predicted mesh. L1 computes the mean absolute distance, while L2 computes the mean squared distance.
- `Normal Consistency (NC)` — range: [0, 1]
  - Cosine similarity between the surface normals of the predicted mesh and the normals of the ground truth point cloud. Values closer to 1 indicate better alignment.
- `Recall@2.5mm (R2.5)` — range: percent
  - Percentage of ground truth points whose nearest predicted mesh surface point lies within a 2.5mm threshold. Higher values indicate better coverage.

## Input / output format

**Input**: Single RGB image of a human face (expressive or neutral).

**Output**: Reconstructed 3D face mesh (vertices and faces) or point cloud.

## Scoring recipe

```python
# 1. Rigidly align prediction to GT pointcloud via landmarks + ICP
aligned_pred = align_icp(prediction_mesh, gt_points)
# 2. Remove non-facial areas from GT using segmentation masks
gt_clean = remove_non_facial(gt_points, segmentation_mask)
# 3. Compute metrics
chamfer_l1 = mean_dist(gt_clean, aligned_pred.surface)
chamfer_l2 = mean_dist_sq(gt_clean, aligned_pred.surface)
nc = cosine_similarity(gt_normals, aligned_pred.normals)
r25 = count(gt_clean, dist_to_surface(aligned_pred) <= 2.5mm) / len(gt_clean)
```

## Common pitfalls

- Failing to rigidly align predictions to ground truth via landmark correspondences and ICP before computing distances.
- Not removing non-facial regions (hair, neck, ears, mouth interior) from ground truth using segmentation masks, which inflates error metrics.
- Evaluating only neutral faces, missing the posed expression reconstruction capability which is a key contribution of the benchmark.

## Evidence (verbatim from paper)

> To measure the performance of a reconstructed posed or neutral 3D face, we follow established practice and first rigidly align the prediction to the ground truth pointcloud via landmark correspondences and ICP. Furthermore, we use segmentation masks[[51]] to remove non-facial areas (hair, neck, ears, and mouth interior) from the ground truth. We then compute three metrics: (i) uni-directional Chamfer distance (L1 and L2) from GT points to the nearest mesh surface, (ii) cosine similarity (NC) of predicted mesh normals and GT pointcloud normals, (iii) Recall thresholded at 2.5mm (R2.5) which is the percentage of GT points whose nearest mesh surface is 2.5mm or closer.

## Citation

```bibtex
@misc{giebenhain2025pixel3dmm,
  title={Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction},
  author={Giebenhain et al. (2025)},
  year={2025},
  note={arXiv:2505.00615}
}
```

- arXiv: 2505.00615

