A synthetic person is a digitally fabricated human being, one whose face, voice, body movements, and sometimes personality are generated entirely or largely by artificial intelligence rather than captured from a single real individual. These creations range from still images of faces that never existed to fully animated, speaking, thinking digital humans that can hold conversations in real time. Building one involves stitching together several AI systems, each handling a different slice of what makes a person seem real: appearance, sound, motion, and behavior.
How a Synthetic Face Gets Made
The visual foundation of most synthetic people starts with generative adversarial networks, or GANs. The basic idea is that two neural networks compete: one generates images while the other tries to spot fakes. Over millions of rounds, the generator gets so good that its output is indistinguishable from photographs of real people. GANs were responsible for the early wave of photorealistic fake faces that circulated widely online starting around 2018, and they remain a core tool in the pipeline.
More recently, diffusion models have entered the picture. These systems learn to gradually remove noise from a random static pattern until a coherent image emerges, and they tend to produce even more realistic and controllable results than GANs alone. Research has shown that diffusion models can enhance frames from video to make them appear more realistic and harder for detection systems to flag, essentially raising the bar on visual quality beyond what earlier GAN architectures achieved on their own.1Knowledge-Based Systems. Diffusion model in modern detection: Advancing Deepfake techniques Both approaches, alongside newer techniques like Neural Radiance Fields, have contributed to increasingly sophisticated synthetic video generation.2arXiv. Deepfake Synthesis vs. Detection: An Uneven Contest
Moving Beyond Flat Images Into Three Dimensions
A still-photo synthetic person is impressive but limited. To create someone who can turn their head, change expressions, and be viewed from any angle, creators need a three-dimensional model. Neural Radiance Fields have become a go-to method for this. The technique works by training a neural network to understand how light interacts with a 3D scene, then using that understanding to render new viewpoints that were never directly captured. Applying this to human avatars means a synthetic person can be rotated, relit, and animated with a level of realism that flat images cannot match.
One persistent challenge is that these models are computationally heavy. Generating every possible viewing angle for a human body requires evaluating enormous numbers of sample points in 3D space, which eats up storage and processing time. Newer lightweight models aim to overcome that bottleneck, making NeRF-based avatars more practical for real-time applications.3Pattern Recognition. LiteNeRFAvatar: A lightweight NeRF with local feature learning for dynamic human avatar For heads and faces specifically, hybrid approaches that combine NeRF’s visual richness with structured facial models produce high-fidelity avatars that can express nuanced emotions while staying true to a consistent identity.4ACM Transactions on Graphics. HAvatar: High-fidelity Head Avatar via Facial Model Conditioned Neural Radiance Field
An even newer rendering technique called 3D Gaussian Splatting has begun competing with NeRF for avatar work. It represents scenes as collections of small 3D blobs rather than volumetric fields, and renders faster as a result. However, it can suffer from geometric drift over long animation sequences, where the synthetic person’s features slowly degrade and lose identity-specific details. Recent work on anchoring these Gaussian models to identity data has shown strong results, maintaining facial detail accuracy even thousands of frames into an animation.5PubMed Central. IA-3DGS Identity-Anchored Dynamic Gaussian Splatting for Long-Term Facial Detail Preservation
Giving a Synthetic Person a Voice
A convincing synthetic person needs to sound human, and voice cloning technology has advanced rapidly. Modern voice synthesis systems can clone someone’s voice from remarkably small amounts of audio. So-called zero-shot voice cloning requires only a few seconds of reference speech to produce new utterances in that person’s vocal style, capturing timbre, pacing, and even emotional inflection. These systems are surveyed as part of a growing body of research into controllable speech synthesis.6IEEE Access. Voice Cloning: A Survey of Zero-Shot and Controllable Speech Synthesis
The voice component does not exist in isolation. It has to be synchronized with the synthetic face’s lip movements, jaw, and even subtle head tilts that real people make while talking. Getting this audiovisual alignment wrong is one of the fastest ways to break the illusion, because humans are extraordinarily sensitive to mismatches between what they see and hear in a face. Research labs have developed systems that learn lip synchronization directly from audiovisual data, training networks on paired video and audio so the resulting avatar’s mouth movements match the generated speech naturally.
For training these visual systems, researchers sometimes build massive synthetic datasets. One approach renders hundreds of thousands of paired image samples under controlled virtual lighting conditions, creating a large library of faces with known attributes that can be used to train relighting and synthesis models.7arXiv. Learning to Relight Portrait Images via a Virtual Light Stage and Synthetic-to-Real Adaptation These synthetic training sets help bridge the gap between laboratory conditions and the messy lighting of real-world footage.
Adding a Mind and Personality
The most ambitious synthetic people are not just visual and auditory constructs. They have a personality layer, something that governs what they say, how they react, and what they remember. This is where large language models enter the picture, the same technology behind modern chatbots, but tuned to behave consistently as a specific person rather than a generic assistant.
A recent research framework describes what it calls a “Human Digital Twin” that integrates a large language model with dynamically updated personal data, enabling it to mirror an individual’s conversational style, memories, and behaviors.8PubMed Central. Towards the “Digital Me”: A vision of authentic Conversational Agents powered by personal Human Digital Twins The idea is that a synthetic version of you would not just look and sound like you; it would remember your preferences, tell your stories, and respond to questions the way you would. Achieving this requires feeding the system a continuous stream of personal data, everything from social media posts and emails to recorded conversations and schedule data.
This personality modeling sits at the heart of what separates a synthetic person from a simple deepfake video. A deepfake swaps one face onto another body in existing footage. A fully realized synthetic person can generate novel responses to novel situations, improvise a conversation, and maintain a coherent persona over time. The distinction matters because it determines what these creations can do: a deepfake replays or manipulates recorded moments, while a synthetic person can be autonomous.
Pulling All the Pieces Together in Real Time
Each component described above, vision, voice, motion, personality, is typically developed by separate research teams using separate models. The frontier challenge is running them all together simultaneously, so a synthetic person can see, hear, think, and respond in real time the way a human does in a video call.
One system, called U-Mind, represents a step toward this unified approach. It jointly models language, speech, motion, and video synthesis within a single interactive loop, so the synthetic person can carry on a multimodal dialogue where its speech, facial expressions, and body language all stay synchronized.9arXiv. U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation The system uses an internal reasoning step before generating output across modalities, which helps ensure that what the avatar says, how it sounds, and what its face does all remain temporally aligned. Getting this wrong, even by a fraction of a second, triggers the kind of uncanny dissonance that makes viewers deeply uncomfortable.
The Uncanny Valley Problem
Almost-but-not-quite-human synthetic people tend to provoke a distinctive negative reaction. This is the “uncanny valley” effect, where entities that are very close to human appearance but slightly off trigger feelings of unease or even revulsion. Research has corroborated this phenomenon experimentally. In one set of studies, participants rated emotional responses to a range of artificial and human faces and also performed threat-perception tasks. The results showed that the uncanny feeling may function as a warning system, alerting people to entities that blur the line between animate and inanimate. Faces that took participants longer to classify as “real” or “unreal” were also the faces associated with the strongest negative emotional responses, suggesting that the discomfort is tied to categorical uncertainty about whether something is alive.10PubMed. Human Perception of Animacy in Light of the Uncanny Valley Phenomenon
For synthetic person creators, this means that a 90%-convincing avatar can actually perform worse than a 70%-convincing one in terms of audience reception. The closer you get to perfect realism, the more any remaining flaw stands out, whether that is a slightly glassy eye, unnatural skin texture, or the wrong micro-expression at the wrong moment. Paradoxically, overtly stylized or cartoon-like virtual characters often avoid this problem entirely because no one expects them to be real.
Virtual Influencers and Entertainment
One of the most visible applications of synthetic people is in social media, where AI-generated virtual influencers already have millions of followers. These characters post selfies, endorse products, and interact with fans despite not being real. Research into how audiences respond to these entities has identified distinct patterns. A qualitative study of Argentine Gen-Z audiences found three reception modes: uncritical acceptance driven by visual appeal, skeptical negotiation based on informational usefulness, and conscious rejection grounded in a demand for “ontological authenticity,” meaning these viewers wanted the people they followed to actually exist.11Global Media Journal México. La empatÃa sintética de los influencers virtuales en Argentina
The concept the researchers describe as “synthetic empathy” is worth noting. It refers to the emotional response audiences have toward an entity that has no inner life. The avatar cannot actually feel anything, but viewers can still develop genuine affective bonds with it, especially when the platform’s features (comment sections, story replies, algorithmic recommendations) encourage a sense of relationship. Whether these bonds are meaningfully different from the parasocial relationships fans have always had with fictional characters remains an open question, but the commercial implications are clear: synthetic influencers are cheaper to manage than human ones, never age, never have scandals, and work around the clock.
Synthetic People in Politics and Public Life
Synthetic persons are also appearing in more consequential contexts. Research into AI-driven political actors has found that some experiments involve AI systems running as candidates or forming political entities, though these vary widely in ambition and seriousness. Some pursue actual electoral goals; others are primarily symbolic or experimental. What they share is a shift away from human-centered political participation toward algorithmically mediated representation, raising questions about what it means for a non-human entity to represent human interests.12PubMed Central. Synthetic political actors: an exploration into emerging AI-driven candidates and parties
These synthetic political actors do not form a single ideological camp. They differ in how much authority is held by humans versus AI, how they justify AI’s political role, and how they organize participation. Most are marked by limited organizational development and interface-centered engagement, more like interacting with a chatbot than attending a rally. The phenomenon is still nascent, but it highlights how synthetic person technology is migrating from entertainment into domains where trust and accountability carry much higher stakes.
Detecting Synthetic People and Proving What Is Real
As the technology for creating synthetic people improves, so does the arms race to detect them. Detection systems face a fundamental asymmetry: creators can train on the detectors’ weaknesses, while detectors are always reacting to the latest generation technique. Several complementary strategies are being developed to address this.
Digital watermarking is one approach, where an invisible signature is embedded into media at the point of creation. If the content is later altered or misattributed, the watermark can flag the manipulation. One proposed framework combines attention-driven watermark embedding with blockchain-based verification, achieving roughly 97% accuracy in identifying salient regions for watermark placement while preserving high visual quality. The system detected all face-swapping attacks in testing and showed strong resilience against common degradations like video recompression and blurring.13PubMed Central. An integrated framework for proactive deepfake mitigation via attention-driven watermarking and blockchain-based authenticity verification
However, watermarking itself can be weaponized. Research has shown that deepfake images embedded with certain types of watermarks can actually fool existing detection models, because the watermark artifacts mask the telltale signs detectors rely on. One study found that a robust detection model could outperform baseline detectors by roughly 10 to 20 percentage points on watermark-contaminated datasets, suggesting that detector architectures need to account for watermarking as a potential evasion strategy, not just a protective measure.14PubMed Central. Robust deepfake detector against deep image watermarking The broader trend in detection research is moving toward transformer-based architectures that can work across multiple media types, attempting to consolidate watermark embedding, forgery detection, and content authentication into integrated systems.15PubMed Central. Multimodal transformer-based watermarking for deepfake detection and digital media authentication
Legal Questions Around Synthetic Identity
The law is struggling to keep pace with synthetic person technology. When an AI system replicates someone’s face or voice, traditional intellectual property frameworks were not designed to handle the resulting disputes. The “right of publicity,” the legal principle that you control commercial use of your own likeness, was written with real humans in mind. Legal scholarship has begun examining whether these frameworks could extend to virtual or AI-generated influencers, analyzing how existing state-level statutes and common law might apply to entities that look and sound like real people but are entirely fabricated.16The Columbia Journal of Law & the Arts. AI Influencers and a Right of Publicity
Voice cloning presents a particularly thorny problem. When an AI replicates a singer’s voice for use on a digital platform, some jurisdictions lack explicit protections for the voice as a distinct legal object. In Indonesia, for example, researchers found no law specifically governing a singer’s voice as a publicity right, though protection could theoretically be constructed through analogies to copyright law, personal data protection statutes, and electronic transaction regulations.17Law Research Review Quarterly. Juridical Review of Singer’s Voice Publicity Rights in AI Use on Digital Media Platform This patchwork approach, cobbling together protections from laws designed for different purposes, is common globally and leaves significant gaps.
Bringing Back the Dead
Perhaps the most emotionally charged application of synthetic person technology is digital resurrection, using AI to recreate the voice, image, and personality of someone who has died. Grieving families can now commission synthetic versions of deceased relatives for conversation or companionship. Entertainment companies have explored resurrecting performers for new content. These applications raise ethical concerns that go beyond standard intellectual property disputes.
Research in neuroethics has framed digital resurrection as a challenge involving mental privacy, cognitive liberty, and the authenticity of AI-generated representations. A proposed governance framework for this space rests on three principles: protecting the mental privacy of the person being represented, ensuring faithful representations of their identity rather than distortions, and preventing exploitation of the deceased or their families. Practical mechanisms suggested include “digital neural wills” that let people specify in advance how their likeness may be used after death, tiered regulation based on the capability level of the technology involved, and structured family decision-making processes.18PubMed. Digital Resurrection and Posthumous Identity: Toward a Cross-Cultural Neurorights Framework
The concept of a digital neural will is particularly interesting because it acknowledges something current law largely ignores: that a person might have preferences about their own synthetic afterlife. Without such instruments, decisions about digital resurrection fall to family members, estate holders, or companies that own relevant data, none of whom necessarily know what the deceased person would have wanted. This gap is likely to become more pressing as the technology becomes cheaper and more accessible, making it possible for anyone with enough photos and audio recordings to build a rough synthetic version of someone they have lost.
How Synthetic People Get Trained
Behind every synthetic person is a training pipeline that consumes enormous amounts of data. For visual models, this might include millions of photographs of real human faces scraped from the internet, or carefully constructed synthetic datasets rendered under controlled conditions. The relighting pipeline mentioned earlier generated about 300,000 paired image samples at 512×512 resolution, stored in high dynamic range format, just to train a model that handles how light falls on faces.7arXiv. Learning to Relight Portrait Images via a Virtual Light Stage and Synthetic-to-Real Adaptation That is one sub-component of one feature of a synthetic person’s visual system.
Voice models require large audio corpora. Personality models need text data reflecting the target person’s communication patterns. Motion models train on motion-capture data or video of real human movement. Each layer of the synthetic person has its own data appetite, and satisfying all of them while respecting privacy and consent is one of the field’s ongoing tensions. Many of the most capable face-generation models were trained on datasets assembled without the knowledge or permission of the people whose faces appeared in them, a practice that has drawn increasing legal and ethical scrutiny.
Where the Gaps Still Are
For all the progress described above, building a fully convincing synthetic person that holds up under sustained real-time interaction remains extremely difficult. The individual components, face generation, voice synthesis, language modeling, motion capture, are each impressive in isolation, but combining them introduces compounding errors. A slightly off lip sync combined with a slightly robotic vocal inflection combined with a slightly delayed facial expression creates a cumulative uncanny effect far worse than any single flaw alone.
Touch and physical interaction remain essentially unsolved. A synthetic person exists only as light on a screen and sound from a speaker. Haptic feedback research is advancing in adjacent fields, but integrating it with the audiovisual pipeline of a synthetic person is largely unexplored territory. Smell, temperature, and the dozens of other subtle cues that contribute to human presence in a shared physical space are not even on the near-term horizon. For now, synthetic people live entirely in screens, making them suitable for video calls, social media, and digital assistants, but fundamentally limited in contexts that demand physical co-presence.