Skip to content

Markerless · phone video in, character video out

AI motion capture from video — no suit, just the finished clip

Film a performance on your phone — a fight move, a dance, a speech — and your character performs it in a finished video. No suit, no markers, no studio. If you need animation data for a rig, a dedicated mocap app is the right tool; this page explains both.

  • No suit or markers
  • Phone video is enough
  • Body, hands and face
  • Finished MP4, not FBX

Updated by the

Made with Kenerate

Quick answer

Updated

What is AI motion capture?

AI motion capture reads a performer's movement from ordinary video, with no suit or markers. Mocap apps turn it into skeleton data (FBX or BVH) for a 3D rig; Kenerate AI's Motion Transfer tool instead puts the performance straight onto your character as a finished MP4, on 12 engines, from takes up to 2 minutes.

Key facts

Hardware
Any phone or camera
Suit or markers
None
Output
Finished MP4 video
FBX / BVH / skeleton
Not exported
Longest take
2 min · SCAIL 2
Face + lips
DreamActor v2, Kling 3.0

Real results

Performances worth capturing

A longboard carve, a sword draw, a pirouette, a little robot's roll and a violinist's bowing — body, props and fine hand work.

These example clips were generated with image to video (model on each card) to show the kind of shots. With Motion Transfer, the performance comes from the take you film.

How to

Capture a performance in five steps

Time needed
Steps
4 steps
  1. Step 1

    Set up the take

    Phone on a tripod, performer fully in frame, plain background, even light. No suit, no markers, no calibration.

  2. Step 2

    Perform and trim

    Record the take, then trim it to the action. Kling takes 3–30 s; SCAIL 2 up to 2 min. MP4, MOV, WEBM, M4V or MKV, around 100 MB.

  3. Step 3

    Choose the character

    A clear JPG or PNG of the character who should perform it, framed like the take.

  4. Step 4

    Pick the solver engine

    Kling 3.0 Pro Motion for hands and faces, DreamActor v2 for facial performance, SCAIL 2 for long takes, LTX-2 Control to re-render the scene.

Kenerate AI Motion Transfer with a full-body photo of an explorer character added as the character image

Open Motion Transfer and upload a clear photo of the character, or choose one from your history.

1 / 5 · Pick a character image

Watch the tour · 1:19

Text, image & reference to video, models, edit, extend & chat

Chapters

Checked 3 October 2026

Video-to-video capture vs mocap apps

Mocap app facts from each product's own page, checked 3 October 2026. Kenerate is our product; the others export animation data, which Kenerate does not.

Video-to-video capture vs mocap apps — checked 3 October 2026
Kenerate Motion TransferRokoko VisionDeepMotion Animate 3DViggle PINOC
What you getFinished MP4 of your character3D animation data3D animation data3D skeletal animation
Export formatsMP4 onlyFBX; BVH on paid tiersFBX, BVH, GLB, MP4FBX or GLB
Needs a 3D rigNo — one character imageYes, to use the dataYes, to use the dataYes, to use the data
Hands and faceHands and face on Kling 3.0 and DreamActor v2Body; fingers not supported at launchFace and hand trackingFingers and wrist rotation in the skeleton
Best forCreators who want the final shotAnimators in Blender, Unreal, UnityGame and avatar pipelinesMixamo-style rigs

If your next step is a 3D scene, a game engine or a rig, use a mocap app. If your next step is posting or editing the video, Motion Transfer skips the rig entirely.

Sources: Kenerate Motion Transfer tool · Rokoko Vision — AI motion capture (exports, free tier) · DeepMotion — Animate 3D · Viggle — PINOC AI motion capture · Autodesk — markerless motion capture

Engines

Which engine captures which part of a performance

In video-to-video motion capture, the engine is the 'solver'. Pick it by what matters most in your take: the whole body, the hands, the face or the whole scene.

  • Kling 3.0 Pro Motion title card — a precise choreographed routine copied from a dancer onto a 3D character
    Detail

    Kuaishou

    Kling 3.0 Pro Motion

    Input
    Image + 3–30 s video
    Output
    Follows the source · sound switch

    Hands, cloth and faces — the most faithful capture of a 3–30 s take, sound kept with the switch

    Open Kling 3.0 Pro Motion
  • DreamActor v2 title card — a fictional adult dancer mid-move with an expressive face
    Face

    ByteDance

    DreamActor v2

    Input
    Image + video up to 30 s · no prompt
    Output
    Follows the source · keeps sound

    Performance capture: body, facial expression and lip movement, no prompt, sound kept

    Open DreamActor v2
  • SCAIL 2 title card — a fictional adult dancer being replaced mid-routine by a different stylised character

    WaveSpeed

    SCAIL 2

    Input
    Image + video up to 2 min
    Output
    480p or 720p · keeps sound

    Long takes up to 2 min; 3D-aware pose holds through turns; sound kept

    Open SCAIL 2
  • Wan 2.2 Animate 2 title card — an open animation model clip of a puppet coming to life

    Alibaba

    Wan 2.2 Animate 2

    Input
    Image + video up to 2 min · motion prompt
    Output
    480p or 720p · silent

    Body, hand gestures and expression on long takes, with a motion prompt

    Open Wan 2.2 Animate 2
  • LTX-2 Control title card — a pose skeleton

    Lightricks

    LTX-2 Control

    Input
    Video up to 20 s (+ optional image)
    Output
    480p, 720p or 1080p · keep, new or mute sound

    Re-render the whole scene from the take's pose, depth or edges, up to 1080p

    Open LTX-2 Control
  • Kenerate Motion title card — a dancer's silhouette whose motion is being copied onto a stylised character

    Kenerate

    Kenerate Motion

    Input
    Image + video (first 20 s used)
    Output
    720p or 1080p · 24/48 fps · sound switch

    Quick 1080p capture of the first 20 s at original, 24 or 48 fps

    Open Kenerate Motion

Every engine returns a finished MP4 with no watermark — never skeleton, FBX, BVH or GLB data. Over-long takes are trimmed on SCAIL, Wan, Kling and Kenerate Motion, and refused on DreamActor v2 and LTX-2 Control.

Tips

Shooting a clean take

Markerless capture is only as good as the footage. These habits come from mocap practice and apply here too.

  • Lock the camera — handheld shake reads as body movement.

  • Fitted clothes that contrast with the background.

  • Whole body in frame, feet included, for full-body moves.

  • Turn the performer's hands toward the camera for gestures.

  • One performer per take; others in frame confuse the solver.

  • Only capture people who agreed to it, onto characters you have rights to.

Example uses

How creators use AI motion Capture

Illustrative examples of typical workflows, not customer reviews.

  • Singer performing in a rehearsal room while a tripod phone records

    Music video creator

    A creator can film the singer's whole performance once and have an animated character perform it as the final music video, without building a rig.

  • Stunt performers rehearsing a fight while a coordinator watches a tablet

    Stunt coordinator

    A stunt coordinator can preview a fight beat on a costumed character from a rehearsal take, then hand the same take to a mocap app if the 3D team needs data.

  • Presenter gesturing to a phone camera in a home office

    Explainer video producer

    A producer can capture a presenter's gestures and expressions and give them to an illustrated host for a course video.

FAQ

AI motion capture questions

Can AI do motion capture from a normal video?

Yes. AI motion capture, also called markerless mocap, reads a performer's body, and often hands and face, from ordinary video shot on a phone or webcam. Mocap apps turn that into 3D animation data; Kenerate's Motion Transfer tool turns it directly into a finished video of your character performing the same take.

Link to this answer

Does Kenerate export FBX, BVH or skeleton data?

No. Every engine in the Motion Transfer tool returns a finished MP4 video, not animation data, so there is nothing to import into Blender, Maya, Unity or Unreal. If you need FBX or BVH for a rig, use a dedicated mocap app such as Rokoko Vision, DeepMotion Animate 3D or Viggle's PINOC, which list those exports on their own pages.

Link to this answer

Who is video-to-video motion capture for?

Creators who want the final clip rather than a rig: social videos, music videos, ads, pitch reels, explainer hosts and VTuber-style clips. You skip modelling, rigging, retargeting and rendering — one character image and one filmed take give you the finished shot. Game and 3D animators who need editable curves should use a mocap app instead.

Link to this answer

Do I need a mocap suit or markers?

No. A phone on a tripod is enough: the engines read the movement from ordinary footage. What helps is what helps any markerless system — a locked camera, the whole body in frame, fitted clothes that contrast with the background, even lighting and one performer at a time.

Link to this answer

Can it capture facial expressions and lip movement?

Yes, on the right engines. DreamActor v2 copies body, facial expression and lip movement onto your character and keeps the take's sound, with no prompt. Kling 3.0 Pro Motion is the most faithful for faces and hands on 3–30 second takes. For full-body takes longer than 30 seconds, SCAIL 2 keeps the body steady through turns.

Link to this answer

How long can a captured take be?

Up to 2 minutes on SCAIL 2, SCAIL, the Wan engines and SteadyDancer; longer takes are trimmed. Kling takes 3–30 seconds, DreamActor v2 up to 30 and LTX-2 Control and Kenerate Motion up to 20. You pay per second of take, so trim to the performance before uploading; failed jobs are refunded automatically.

Link to this answer

Can I capture someone else's performance?

Only with their permission. Film yourself or performers who agreed to it, and apply the performance only to characters you own or have rights to — never celebrities. Captured performances are convincing, so label the result as AI-made when you publish it.

Link to this answer

Two kinds of AI motion capture

Data mocap tracks the performer's joints and exports a skeleton animation — FBX, BVH or GLB — that you retarget onto a rigged 3D character and render yourself. Video-to-video capture, which is what Kenerate does, skips the skeleton you can edit: the engine reads the take and renders your character performing it, lighting and all, as a video. Both start from the same thing: a clear take of a real performer.

Explore

Every Kenerate video tool

Generators, tools, models and guides — all in one studio.

See more of Kenerate AI in your Google resultsAdd kenerateai.com as a preferred source — Google shows more from your preferred sources in Top Stories and AI answers.

Turn a phone take into a finished shot

No suit, no markers, no rig. One take, one character image, a finished MP4.

Tried it? Tell us how it went