LAB 001 · Computer vision / Foot interaction
Markerless Foot Tracking
From a color-marker baseline to body pose and a foot-specific six-keypoint research path.
- Question
- Can one phone camera turn ordinary foot movement into stable interaction without colored markers?
- Status
- Promising research
- Year
- 2026
- Related work
- Phi Step ↗
—
Current investigation
- Camera frame
- Foot6 keypoints
- Weighted foot anchors
- Temporal tracking
- Player-space calibration
- Step events
01
Question
Phi Step needs repeatable left and right foot positions for a virtual five-zone mat. Could that input work with ordinary shoes and lower-body framing rather than colored socks or markers?
02
Approaches
The OpenCV HSV and contour path used sampled colors to validate the camera-to-zone-to-event loop. It made the interaction testable, but required visible colored markers and careful separation of the two feet.
MediaPipe Pose Landmarker removed the marker requirement. In physical comparison, it still needed nearly full-body framing; waist-to-feet framing was unreliable.
03
Current investigation
A foot-specific research model produces left and right big toe, small toe, and heel points from a fixed lower-body camera frame. Accepted points form weighted anchors; temporal tracking handles brief uncertainty before player-space calibration and the existing Step Event Engine classify entries.
04
Evidence
Pretrained RTMW was compared on real Phi Step standing and movement recordings across different frame crops. Each run produced annotated video, frame-level keypoint CSV, and summary and timing output. The model-family result was promising, but availability and visual continuity were not ground-truth accuracy.
A later six-point Core ML research checkpoint and supplied iPhone recordings support a limited lower-body vertical slice with ordinary footwear and step-event judgement. They do not establish general reliability or a public release.
05
What remains open
Crossing or close feet, longer occlusion, frame edges, anatomical identity across diverse people, and a complete physical rhythm-game session still need stronger evidence. The next useful test is independent, representative evaluation across those cases.