← All experiments

LAB 001 · Computer vision / Foot interaction

Markerless Foot Tracking

From a color-marker baseline to body pose and a foot-specific six-keypoint research path.

Question
Can one phone camera turn ordinary foot movement into stable interaction without colored markers?
Status
Promising research
Year
2026
Related work
Phi Step ↗

—

Current investigation

  1. Camera frame
  2. Foot6 keypoints
  3. Weighted foot anchors
  4. Temporal tracking
  5. Player-space calibration
  6. Step events

01

Question

Phi Step needs repeatable left and right foot positions for a virtual five-zone mat. Could that input work with ordinary shoes and lower-body framing rather than colored socks or markers?

02

Approaches

The OpenCV HSV and contour path used sampled colors to validate the camera-to-zone-to-event loop. It made the interaction testable, but required visible colored markers and careful separation of the two feet.

MediaPipe Pose Landmarker removed the marker requirement. In physical comparison, it still needed nearly full-body framing; waist-to-feet framing was unreliable.

03

Current investigation

A foot-specific research model produces left and right big toe, small toe, and heel points from a fixed lower-body camera frame. Accepted points form weighted anchors; temporal tracking handles brief uncertainty before player-space calibration and the existing Step Event Engine classify entries.

04

Evidence

Pretrained RTMW was compared on real Phi Step standing and movement recordings across different frame crops. Each run produced annotated video, frame-level keypoint CSV, and summary and timing output. The model-family result was promising, but availability and visual continuity were not ground-truth accuracy.

A later six-point Core ML research checkpoint and supplied iPhone recordings support a limited lower-body vertical slice with ordinary footwear and step-event judgement. They do not establish general reliability or a public release.

05

What remains open

Crossing or close feet, longer occlusion, frame edges, anatomical identity across diverse people, and a complete physical rhythm-game session still need stronger evidence. The next useful test is independent, representative evaluation across those cases.