Crowds and Multiple Characters in AI Video: Why They Break and How to Direct Them

Crowds and Multiple Characters in AI Video: Why They Break and How to Direct Them

October 07, 2026•6 min read

Put more than a couple of characters in a Seedance shot and two different things start to break. Either the people mush together into smears of colour that do not read as individuals, or the model puts the right people in the wrong places, sitting on the wrong side of the aisle, facing the wrong way, standing where nobody told them to stand. These are two separate problems with two separate fixes, and mixing them up wastes a lot of generations.

Mush happens when you bake the crowd into the reference

I ran into this directly on a ski race shot where a crowd behind the starting gate had been baked directly into the reference image. From that distance the people came out as mushed splashes of colour rather than anything that read as a person. My fix, and the rule I use for every crowd shot now, never generate the crowd into your reference image. Give Seedance the empty environment instead, empty arena, empty stand, empty street, and let the crowd itself be part of the prompt, not the picture.

The reasoning is worth understanding, not just copying. When Seedance has an empty plate to work from, it imagines the crowd's pixels fresh, in whatever form fits the video generation best. When you hand it a crowd that is already baked into a still image, it tries to push those exact pixels into motion, and pixels that were never built to move do not hold up as people once they do.

There is a second fix worth stacking on top of that, a depth of field cheat. Keep four or five people sharp in the foreground and let everyone behind them fall out of focus, the same way portrait mode on a phone does it. Viewers read a crowd as real from the sharp faces at the front. What is happening in the blur behind them barely registers, and it does not need to hold up to scrutiny it will never get.

One limit is worth knowing so you stop fighting it. Very high aerial shots with tiny, distant figures remain a genuine model limit right now, not a prompting problem. If you need that shot, plan for a different angle rather than spending the budget trying to prompt your way past it.

Misplacement is almost always the prompt, not the model

The second failure mode showed up clearly for me when I was trying to seat a character on the window seat of a plane and kept getting the aisle seat instead, generation after generation. My first move was not to fix the model, it was to read the prompt Claude had written. It named both seats, aisle seat and window seat, in the same sentence, describing what was next to the character rather than only what the character was doing. My rule, strip every word that is not the instruction itself. Telling the model what is beside the character, instead of only where the character is, gives it two things to satisfy and it will sometimes pick the wrong one.

When stripping the prompt is not enough, the next move is a position reference, a separate image showing the characters already placed where you want them in the space, built quickly in an image model and tagged in the prompt as a position reference, not a location reference and not a starting frame. That distinction matters. A location reference tells the model what the room looks like. A position reference tells it where the bodies go inside that room, and Seedance reads the two differently.

For placing background extras without spending real time on each one, my method is to draw an inner circle and an outer circle directly onto the location image and label them in red text, inner circle and outer circle, then reference which characters belong in which ring in the prompt. The trick that makes this work, annotations drawn onto a reference image are read by Seedance as instructions, they are not rendered into the output. You are writing directions on the photo, not decorating it.

Cloning one person into a crowd of themselves

A specific version of the positioning problem is the same trick behind those clones-of-myself videos that circulate online, five or six versions of one person acting independently in a single shot. The mechanic is simpler than it looks. Upload the same character sheet multiple times, tag each upload as a different named character, Bob, Tom, John, whatever names keep them distinct in the prompt, add the location reference, then describe what each named character is doing second by second inside one continuous shot. The model does not know or care that Bob, Tom and John are the same face. It just needs each name tied to an action it can follow.

The same mechanic works with entirely different people. Swap any of those named references for a different character sheet and you get an ensemble cast placed and acting independently in one generation, which is really the same problem as the aisle seat, just with more seats to fill correctly.

Diagnose before you spam generate

If the people in your shot look like smears, the fix lives in the reference, give Seedance an empty plate and let it build the crowd, or fake depth with a sharp foreground and a blurred back row. If the people are in the wrong spot, the fix lives in the prompt, strip the fluff first, then escalate to a position reference or labelled circles if stripping the prompt is not enough. Treat those as two different diagnostics, not one, and you will stop wasting generations solving the wrong problem.

Common questions

Why do crowds in AI video look mushy or blurred together?

Because the crowd was baked into the reference image. Give Seedance an empty environment instead and prompt the crowd separately, so it generates fresh pixels rather than trying to animate a still that was never built to move.

How do I control exactly where each character sits or stands in an AI video?

Check the prompt first, most misplacement is unnecessary detail confusing the model, strip anything that is not a direct instruction. If that does not fix it, build a position reference image showing the characters already placed, and tag it as a position reference, not a location.

How do I direct large crowds of extras without a reference for each one?

Draw labelled circles on the location image, an inner circle for main players and an outer circle for background extras, and reference which characters belong in which ring in the prompt. Annotations on a reference image are read as instructions and are not rendered into the output.

How do I put multiple clones of the same person in one AI video?

Upload the same character sheet multiple times, tag each one as a different named character, add a location reference, then describe each named character's action second by second in the prompt.

Want the full method and the community that runs it every week? Join GenHQ.

Rourke Sefton-Minns

Rourke Sefton-Minns

AI creative educator and founder of GenHQ, the paid community teaching the AI video, image and film workflows brands pay for, plus how to price the work and land clients.

Back to Blog