I created an interactive digital avatar of myself — and you can talk to it
A journalist visited Synthesia’s New York office and had the company create both personal and interactive digital avatars of them, trained specifically to discuss a single article about venture-backed startup fraud. The interactive avatar was built from photos and a two-minute voice recording and runs on a pipeline of voice-to-text, language, text-to-voice and video models; Synthesia also offers enterprise products like Roleplay Sessions and an API for custom avatar applications.

Why It Matters
The piece highlights how quickly realistic, controllable digital likenesses can be produced and integrated into workflows — a development that raises questions about trust, authenticity, and how AI avatars might change roles in PR, training, and journalism. It also shows enterprises can choose underlying models and hosting, signaling broader commercial adoption and customization of avatar technology.
Key Facts
- Company: Synthesia
- Person mentioned: Alexandru Voica, head of corporate affairs at Synthesia
- Valuation: $4 billion (earlier this year)
- ARR milestone: Crossed $100 million in ARR last year (reported by Synthesia)
- Office visited: Synthesia's new office in New York (company originally based in the U.K.)
A reporter accepted an offer from video-generation startup Synthesia to have a personal and interactive digital avatar created at the company’s new New York office. The session involved numerous photographs and a two-minute voice recording; Synthesia produced both a scripted “personal” avatar and an interactive version trained specifically on the reporter’s article about why venture-backed startups commit more fraud than non-VC-backed startups. The interactive avatar was limited to answering questions about that story.
Synthesia positions the technology for enterprises with several product lines: a classic video-creation and distribution platform where avatars read typed scripts; an agentic product called Roleplay Sessions for interactive training and scoring; and an API platform that lets customers combine Synthesia’s video and voice models with other technologies. The company said it reached a $4 billion valuation earlier in the year and reported crossing $100 million in annual recurring revenue the previous year.
Technically, the interactive avatar uses a pipeline of models: voice-to-text to transcribe user speech, an agentic language model to interpret and act on the text, text-to-voice to generate replies, and Synthesia’s video model to animate the avatar. Synthesia uses its own video and voice models by default but allows customers to select alternatives from other labs such as Cartesia, ElevenLabs, Google, or OpenAI, and to choose where avatars are hosted.
The reporter said Synthesia’s team produced the avatars within a couple of days. Reactions from friends were mixed: some found the avatars interesting and creepy, others thought the interactive likeness and voice were less accurate than the scripted personal avatar. The interactive model behaved deterministically — redirecting questions outside its training (for example, personal queries) back to the article topic.
The experience prompted reflection on possible impacts in journalism and other fields. The reporter observed that while avatars could be useful for tasks like training or answering routine queries, trust remains a core element of journalism that may not transfer to AI. At the same time, they noted potential practical appeal outside newsrooms, where a persistent digital version of someone could handle inquiries during absences like holidays.
Keep Reading

Can Cloudflare CEO Matthew Prince save the web from AI?
China is playing a different game when it comes to AI

Can ‘eSUV’ e-bikes really go from trail to town?
