How Synthetic Media will change Hollywood? with AI, Digital Humans, Voice Cloning + Synthesia Demo

by Sander Saar

The host presents an AI-generated clone of himself and introduces the concept, history, and evolution of synthetic media across consumer apps and film production. Along the way it covers Why Video and Why Now, Creating Synthetic Media with Synthesia & Descript and Use Cases: VTubers, Virtual Influencers, and AI Companions, and more. The host looks ahead to autonomous digital humans and how tools like GPT-3, Descript, and Synthesia will enable making films right from a laptop.

For agents: machine-readable ARVP doc at /public/creators/sander-saar/videos/video_qlBrh60bGlY/arvp.json · Browse this video library

Key findings and exact moments

  1. 0:00Introduction to Synthetic Media

  2. 3:03.55Why Video and Why Now?

  3. 5:49.24Creating Synthetic Media with Synthesia & Descript

  4. 10:20.63Use Cases: VTubers, Virtual Influencers, and AI Companions

  5. 12:56.46Risks, Ethics, and Fingerprinting

  6. 15:28.27The Future: Autonomous Digital Humans

Chapters

  1. 0:00 Introduction to Synthetic Media

    The host presents an AI-generated clone of himself and introduces the concept, history, and evolution of synthetic media across consumer apps and film production.

    On screen: The presenter gestures and speaks while dynamic infographic slides illustrate the evolution of creative media production tools behind him.

  2. 3:03 Why Video and Why Now?

    An exploration of consumer video consumption habits, business conversion metrics, and recent breakthroughs in speech and image machine learning models.

    On screen: A man in a patterned shirt presents in front of changing digital graphics showing AI benchmarks, facial recognition, and Unreal Engine's MetaHuman interface.

  3. 5:49 Creating Synthetic Media with Synthesia & Descript

    A walkthrough demonstrating how to combine Descript voice cloning with Synthesia avatars to rapidly generate multilingual video content at low cost.

    On screen: A man addresses the camera on the left side while various slides about the benefits and automation workflow of synthetic media appear beside him.

  4. 10:20 Use Cases: VTubers, Virtual Influencers, and AI Companions

    A showcase of current synthetic media applications ranging from VTuber avatars and Instagram influencers to interactive AI companions and corporate comms.

    On screen: The presenter stands against a white studio background on the left while animated graphics and screen recordings demonstrating virtual influencers, Replika, and digital humans appear on the right.

  5. 12:56 Risks, Ethics, and Fingerprinting

    A discussion on the ethical dilemmas posed by synthetic media, public awareness, and the technological need for metadata fingerprinting and attribution.

    On screen: Victor Riparbelli speaks in front of evolving presentation graphics illustrating visual effects costs, fingerprinting, and recognition tools.

  6. 15:28 The Future: Autonomous Digital Humans

    The host looks ahead to autonomous digital humans and how tools like GPT-3, Descript, and Synthesia will enable making films right from a laptop.

    On screen: A presenter speaks in front of a white studio backdrop while digital graphics illustrating AI video creation tools appear on screen.

Visual moment index

What the footage shows, shot by shot. It is derived from the video itself, independent of the commentary. Timestamps link to the exact second.

0:00–0:18 An AI digital avatar of a man introduces the video before the real man appears beside him against a solid white studio background.

0:18–0:24 Two cloned versions of a man in matching patterned shirts stand side by side against a white studio background, speaking to the camera. On screen: POW SMASH

0:24–0:26 A man in a blue patterned shirt stands on the left side of a white studio set, pointing to his right while speaking. On screen: SMASH

0:26–0:41 A man in a patterned blue shirt gestures and speaks directly to the camera against a white studio background. On screen: SYNTHETIC MEDIA Deep Fakes

0:42–0:51 A synthetic deepfake video of Kim Kardashian speaking directly to the camera in a room.

0:51–0:53 The host speaks to the audience with a deepfake video clip of Barack Obama playing on screen behind him.

0:54–1:02 A deepfake video of Barack Obama in a suit and tie speaking in front of an American flag.

1:02–1:07 The host speaks and gestures dynamically beside a white digital graphic in a studio setting. On screen: Any media that is created by AI

1:07–1:12 The presenter stands on the left gesturing towards a large text prompt on a clean white background. On screen: Any media that is created by AI SMASH!

1:12–1:16 The presenter speaks beside a vertical smartphone recording demonstrating facial AR filters. On screen: 5:20 Fire Frame

1:16–1:28 The presenter gestures while speaking beside an on-screen smartphone recording showing various augmented reality face filters in action. On screen: 5:20 Hat and Slim Mustache Cat in Car

1:28–1:33 A presenter speaks in the foreground against a screen displaying 3D CGI facial modeling and cinematic renders of Thanos. On screen: INTRODUCING MASQUERADE

1:33–1:37 The presenter gestures enthusiastically while the background screen demonstrates facial motion capture rigs and 3D modeling interfaces. On screen: Blend_HardLight Blend_Lighter Blend_LinearBurn Blend_LinearDodge Blend_LinearLight Blend_Overlay Blend_PinLight Blend_Screen Blend_SoftLight Blend_Scratchlines…

1:38–1:41 The presenter gestures toward the camera as text slides highlight performance benefits on the screen behind him. On screen: FEATURE QUALITY

1:41–1:42 A presenter stands beside a screen demonstrating motion capture facial tracking alongside CGI rendering. On screen: ADVANCED

1:42–1:46 The presenter gestures while a video screen behind him displays various synthetic face-capture and lighting techniques from Digital Domain. On screen: MACHINE LEARNING REDUCES ARTIST TIME

1:46–2:01 A man gestures expressively beside a digital presentation comparing consumer Snapchat filters to professional VFX in Gemini Man. On screen: CONSUMER PROFESSIONAL weta DIGITAL WILL SMITH GEMINI MAN

2:01–3:08 The presenter gestures and speaks while dynamic infographic slides illustrate the evolution of creative media production tools behind him. On screen: CONSUMER Lens Studio ark AR TikTok PROFESSIONAL UNREAL ENGINE unity snowdrop FROSTBITE Analog Digital AI Roland Instruments Beats Voices Pitch corrections WHY…

3:08–3:38 A man gestures while presenting alongside an infographic slide highlighting video consumption and search engine statistics. On screen: WHY VIDEO? 68% consumers prefer watching videos to learn about new products or services vs articles, infographics, ebooks, ations YouTube No 2 Search engine…

3:38–3:45 An animated graphic displays a dynamic email template sending a personalized video greeting to Sebastian. On screen: WATCH VIDEO If you have any questions feel free to hit reply! - Pranay Prakash Send > Sebastian Woodsworth To: Sebastian Woodsworth Subject: Hello, Sebastian!…

3:46–3:48 The dynamic email template updates to demonstrate another recipient, showing a personalized video for David. On screen: David Edmonson To: David Edmonson Subject: Hello, David! Hey David! I just recorded this short video for you as a quick introduction. WATCH VIDEO If you have…

3:50–3:54 An animated graphic displays an email inbox scrolling on a mobile phone interface over a pink and orange gradient background. On screen: Pranay's Inbox Jason Argonaut Re: Hello, Jason! Wow! That was amazing! Fred Arminsen Re: Hello, Fred! Pranay, good to hear about you my friend... Mery Doe Re:…

3:54–3:59 A man gestures and presents in front of a slide featuring customer statistics and a bar chart comparing repurchase rates. On screen: Customers were 87% more likely to buy from us again when they received the personalized video from me. 3.75% 7.02% Re-purchase rate Didn't receive personalized…

4:00–4:18 The presenter speaks to the camera while gesturing, standing beside an informational graphic and bar chart about a Peel customer case study. On screen: Customers were 87% more likely to buy from us again when they received the personalized video from me. PEEL 3.75% 7.02% Re-purchase rate Didn't receive…

4:18–4:39 A presenter speaks in front of evolving infographic slides detailing case studies, cost savings, and machine learning accuracy trends. On screen: Customers were 87% more likely to buy from us again when they received the personalized video from me. PEEL 3.75% 7.02% Re-purchase rate Didn’t receive…

4:39–5:02 A man in a patterned shirt presents slides explaining Google Machine Learning word accuracy and Google WaveNet voice generation. On screen: WHY NOW? Google Machine Learning Word Accuracy, 2013-2017 100% 90% 80% 70% 2013 2014 2015 2016 2017 Google ML Human accuracy 80+ languages & variants 20sec…

5:02–5:07 Victor Riparbelli gestures as he speaks to the camera, with a graphic diagram illustrating the Turing Test displayed behind him in the virtual studio. On screen: 80+ languages & variants 20sec audio in 1 sec 2% human speech quality Google Wavenet How natural does it sound? 4.21 4.55 WaveNet Human Speech PASSING THE…

5:07–5:48 A man in a patterned shirt presents in front of changing digital graphics showing AI benchmarks, facial recognition, and Unreal Engine's MetaHuman interface. On screen: PASSING THE TURING TEST A B C Cloud AutoML Vision Facial recognition, emotional state, other attributes metahuman_014 FACE Skin Makeup Teeth HAIR Head Beard…

5:48–6:23 A man addresses the camera on the left side while various slides about the benefits and automation workflow of synthetic media appear beside him. On screen: BENEFITS DEMOCRATIZATION Accessibility Lower cost Simplicity SCALE AUTOMATION BUY shopify 1 ConvertKit pier Synthesia zapier

6:23–6:30 A man gestures and speaks beside a diagram demonstrating scale and automation workflow with various software integrations. On screen: SCALE AUTOMATION onvertKit apier nthesia ConvertKit zapier Synthesia BUY shopify

6:30–7:00 A presenter gestures on the left while slide graphics illustrate automation workflows, speed metrics, and software comparisons on the right. On screen: SCALE AUTOMATION ConvertKit zapier Synthesia BUY shopify SPEED Time to create 10min content 600 30 Human AI SIMPLICITY descript 30min Audio 5 x 3min Video…

7:00–7:17 A presenter stands on the left side of the frame speaking and gesturing beside a comparison graphic displaying Descript and Synthesia details. On screen: SCALE SIMPLICITY descript Synthesia 30min Audio 5 x 3min Video Sander 1

7:17–7:18 A medium shot of a man speaking next to a presentation graphic comparing Descript and Synthesia under scale and simplicity headers. On screen: SCALE SIMPLICITY descript 30min Audio Synthesia 5 x 3min Video Sander 1

7:18–7:32 The presenter is positioned in the lower-left corner as a software interface of Descript is demonstrated on screen. On screen: Audio Growth Intro Greeting Video intro Add new... 0:00:13.63 Share Publish Sander (2) Hey. I'm Sander and today I'm really excited to explore with you the…

7:32–7:34 A man in a patterned shirt speaks in the corner overlay while demonstrating an audio export dialog on screen. On screen: SS Share Publish 0:00:13.63 Video intro Audio Growth Intro Greeting Add new... Project access Export EXPORT > AUDIO Exporting: Current Composition Format:…

7:34–7:39 The screen displays the Synthesia dashboard as the man continues talking from the bottom left corner. On screen: Audio Growth Intro Greeting Video intro + Add new... 0:00:13.63 Share SS Publish Sander (2) Hey. I'm Sander and today I'm really excited to explore with you…

7:39–7:41 The man in the patterned shirt remains positioned in the lower corner against a blank light grey background. On screen: Synthesia Home Videos Templates Avatars Create new video Import PowerPoint OUT SIDE THE BOX Agency follow-up TEAM UPDATE Team Update BOARD MEETING Board…

7:43–7:46 A screencast displays the Synthesia video creation interface with an avatar preview, accompanied by a picture-in-picture view of the presenter in the lower left. On screen: Type your video title here Cancel Generate video 1 Add slide Select template Featured templates My templates OUT- SIDE THE BOX TEAM UPDATE Agency follow-up…

7:46–7:50 The presenter continues demonstrating the Synthesia platform UI while browsing different presentation slide templates. On screen: Type your video title here All changes saved Discard draft Generate video TEAM UPDATE LEADERSHIP MEETING Add slide Select template Featured templates My…

7:50–7:57 The presenter demonstrates template selection and switches to the background selection menu in the Synthesia platform interface. On screen: Type your video title here All changes saved Discard draft Generate video TEAM UPDATE LEADERSHIP MEETING Duplicate Delete Select template Templates Featured…

7:57–8:13 A presenter appears in a bottom-left overlay while demonstrating text, background, and avatar customization features in Synthesia. On screen: Type your video title here All changes saved Discard draft Generate video 1 Add slide Select background Colors Images Videos Uploads Search uploads Upload…

8:13–8:18 A presenter speaks in the lower-left corner over a full-screen screencast of the Synthesia interface demonstrating voice selection. On screen: Type your video title here All changes saved Discard draft Generate video Add slide Select avatar, size and alignment Search avatars Nina Paul QuHarrison…

8:19–8:26 The host in picture-in-picture explains custom audio file uploading on the Synthesia dashboard as the language selector is shown. On screen: Type your video title here All changes saved Discard draft Generate video 1 Add slide Type your script Upload your voice To achieve the best results, please…

8:26–8:28 A screen recording of the Synthesia interface showing a video generation pop-up with the host visible in the lower corner. On screen: Type your video title here All changes saved Discard draft Generate video Add slide Select avatar, size and alignment Search avatars Nina Paul QuHarrison…

8:28–8:33 The host gestures in the lower left corner against the Synthesia dashboard displaying template and project options. On screen: Synthesia Home Videos Templates Avatars Create new video Import PowerPoint OUTSIDE THE BOX Agency follow-up TEAM UPDATE Team Update BOARD MEETING Board meeting…

8:33–8:36 The host speaks and gestures in front of a digital infographic comparing analog, digital, and AI production costs. On screen: SCALE COST Analog Digital AI $200,000 $2,000 $20

8:36–9:05 A presenter gestures while explaining cost differences and showing an interactive Synthesia video demo featuring Lionel Messi. On screen: SCALE COST Analog Digital AI $200,000 $2,000 $20 LEAGUE CHAMPIONS Lay's PERSONAL STEP 1 STEP 2 STEP 3 STEP 4 Sander Nice to meet you, Sander! So, what's your…

9:05–9:13 A man in a patterned shirt presents an on-screen graphic demo of a personalized video generator featuring Lionel Messi. On screen: PERSONAL BENEFITS - PERSONALIZED STEP 1 STEP 2 STEP 3 STEP 4 Watch the game Bueno! Finally, when do you want to do this? MESI01 Online ON WEDNESDAY ON TUESDAY…

9:13–9:44 The host speaks to the camera while a video insert of an AI Lionel Messi plays next to him, which later switches to a graphic showing various supported languages. On screen: LAY'S MESSI MESSAGES Powered by Synthesia LANGUAGES ACCESSIBILITY REACH Arabic - Default Bengali - Original Bulgarian - Natural Catalan - Natural Chinese (CN)…

9:45–9:49 The host presents to the camera with an inset video showing David Beckham speaking on screen. On screen: LANGUAGES voice petition. KISWAHILI

9:50–10:05 David Beckham sits at a dining table speaking directly to the camera about malaria as multilingual audio and lip-sync are demonstrated. On screen: isn't just SPANISH half ARABIC every

10:05–10:40 The presenter gestures and speaks in front of evolving background graphics featuring the Synthesia website, use case title cards, and a VTuber streaming demonstration. On screen: Synthesia Features Products Use Cases Resources Company Log in Create account Goodbye cameras, microphones and actors! professional AI videos from text in 50+…

10:41–10:45 A split screen displays an animated 3D VTuber avatar on the left synchronized with the streamer on the right wearing a head-mounted facial tracking rig. On screen: HI-TECH ELF JACUZZI STREAM WITH BIKINIS AND PHYSICS 0 6223 10000 Ends in 7 days not do t see her s b 捕捉我面部的攝影機在這, 用Iphone X xser

10:45–11:05 A split screen displays a stylized 3D avatar on the left mirroring real-time facial expressions and finger movements of a streamer wearing a facial capture helmet and tracking gloves on the right. On screen: HI-TECH ELF JACUZZI STREAM WITH BIKINIS AND PHYSICS 6223 10000 Ends in 7 days ni onl th is bug 我的手指連接著我的手套 ooaaahh

11:05–11:22 A man gestures expressively while standing next to a screen showing Memoji avatars. On screen: AVATARS Desi Perkins Patrick Starrr Memoji

11:22–12:18 The presenter stands against a white studio background on the left while animated graphics and screen recordings demonstrating virtual influencers, Replika, and digital humans appear on the right. On screen: VIRTUAL INFLUENCERS @lilmiquela 3:57 LEARNING & DEVELOPMENT Working from home

12:18–12:20 A man in a patterned shirt speaks toward the camera as synthetic video examples for learning and development appear on screen behind him. On screen: LEARNING & DEVELOPMENT Synthesic Working from home

12:20–12:46 A presenter gestures while discussing synthetic media use cases alongside changing on-screen graphic examples and mobile screen recordings. On screen: CORPORATE COMMS Dan Danson UX Designer Maria Resolow Customer Succes Andreas Candor Front-end developer CREATIVE EXPRESSION Reface App made with reface app

12:46–13:20 The presenter discusses the ethical risks of AI media while relevant video clips and informational slides display beside him. On screen: made with reface app I BRING IT SMASH! RISKS Ethics Content Authenticity Initiative 1 People first. Always. As a company pioneering this new kind of media…

13:20–13:36 A man addresses the camera on the left while a slide presentation on AI ethics and public knowledge displays beside him. On screen: Ethics Content Authenticity Initiative 1 People first. Always. As a company pioneering this new kind of media we're aware of responsibility we have. It is…

13:36–13:47 A man in a patterned shirt gestures as he presents slides comparing historical costumes to modern social media filters. On screen: Education and public knowledge 6th century BC Costumes 21st century Creative & Fun Nature Magic Superpowers TikTok & Instagram filters

13:47–14:12 A man in a patterned shirt gestures while speaking in front of comparative slides displaying costumes vs filters and human vs computer-generated emails. On screen: 6th century BC Costumes Creative & Fun Nature Magic Superpowers TikTok & Instagram filters 21st century 1970s 99% written humans 2000s by computers

14:12–14:27 A man speaks facing the camera in a white studio setting as digital graphics comparing technologies appear on a screen behind him. On screen: 1970s 2000s File Edit View Mail Window Help Compose Reply Reply All Forward Move Delete Trash Print Spam Next Inbox Private Folders Deleted mail Sent mail From…

14:27–15:10 Victor Riparbelli speaks in front of evolving presentation graphics illustrating visual effects costs, fingerprinting, and recognition tools. On screen: WETA DIGITAL WILL SMITH GEMINI MAN Fingerprinting Tools Transparency SHAZAM YouTube Content ID Music in this video Learn more Listen ad-free with YouTube…

15:10–15:40 A presenter gestures while speaking in front of graphic presentations displaying media ownership tools and the future progression of AI personas. On screen: Tools Transparency SHAZAM YouTube Content ID Music in this video Learn more Listen ad-free with YouTube Premium Song Happier Than Ever Artist Billie Eilish…

15:40–15:49 A man in a patterned shirt presents slides explaining synthetic likeness, followed by a video demonstration graphic. On screen: LOOKS LIKE Make anyone say anything Digital Domain’s Charlatan

15:49–15:56 A man speaks directly to the camera against a plain black background.

15:56–16:03 A split-screen comparison shows two versions of the speaker talking side-by-side against a black background.

16:03–16:09 A man against a black background speaks as his face digitally morphs seamlessly into another man's face.

16:09–16:25 A man speaks and gestures next to a presentation slide that transitions into a demonstration video of Google LaMDA. On screen: ACTS LIKE Make things come alive Google LaMDA

16:25–16:27 A presenter gestures while speaking in front of an on-screen visual displaying a Google LaMDA demonstration of planet Pluto. On screen: ACTS LIKE Google LaMDA

16:27–16:34 The presenter continues gesturing with open hands in front of the Google LaMDA presentation slide showing Pluto. On screen: ACTS LIKE Google LaMDA

16:35–16:44 Sundar Pichai stands outdoors on stage next to a large digital display showing an image of Pluto before an interactive text box appears. On screen: I'm so curious about you

16:44–17:09 A split screen displays a rendered model of Pluto on the left alongside a live text transcript of a conversation with Google LaMDA on the right. On screen: I'm so curious about you LaMDA I sense your excitement. Ask me anything. tell me what I would see if I visited You would get to see a massive canyon, some…

17:09–17:34 A presenter speaks enthusiastically in front of a digital slide displaying text and video conference footage. On screen: THINKS LIKE Autonomous Digital Humans

17:36–17:37 A split-screen video conference displays six panelists attending a virtual discussion. On screen: AK

17:37–17:44 A photorealistic digital human replica of Doug Roble speaks inside a virtual loft apartment setting.

17:47–17:48 A high-speed CGI light cycle race and clash sequence plays out from Tron: Legacy.

17:49–17:50 Belle and the Beast dance together in an ornate ballroom surrounded by chandeliers.

17:50–17:51 A dark visual effects shot showing two characters tumbling inside an aircraft cargo hold.

17:51–17:52 A CGI wireframe breakdown transitions to the final shot of two characters floating in an airplane interior.

17:52–17:54 An extreme close-up focuses on the green eyes of a digitally reconstructed woman's face.

17:54–17:55 An elderly-looking, thin man flexes his arms and smiles in front of a bathroom mirror. On screen: Richards Organica

17:55–18:44 A presenter speaks in front of a white studio backdrop while digital graphics illustrating AI video creation tools appear on screen. On screen: THINKS LIKE YouTube video without writing & filming Writer OpenAI GPT-3 Sound descript reSpeeCHER Camera & Editing Synthesia Make Hollywood films with a laptop

18:44–18:45 A medium close-up of a man speaking directly into the camera against a bright white studio backdrop.

18:46–18:55 Two cloned versions of the man stand side by side in a white studio while speaking. On screen: POW SMASH

18:57–19:05 An end screen layout features side-by-side frames of the host and his AI avatar speaking and waving. On screen: WATCH NEXT SUBSCRIBE

Transcript

0:00 Speaker A: Hey, I'm Sander, and today I'm really excited to explore with you the possibilities of synthetic. Speaker B: No, you're not, Sander. Speaker B: I don't think you're going to be exploring any opportunities here. Speaker A: Yeah, well, I look like you, I can speak like you, and I can probably even write better than you. Speaker A: So maybe just let me take this one. Speaker B: No, you're not going to be taking this one. Speaker B: I'm going to be taking this one. Speaker B: But maybe next time. Speaker A: Okay, but don't forget to like this video and subscribe. Speaker B: Yeah, all of this that you just saw was fully, completely computer generated.

0:29 Speaker B: The visual and the audio and the voice as well. Speaker B: I'm really excited to explore with you synthetic media. Speaker B: And when I say synthetic media, when people first come to think about, is actually deep fakes. Speaker B: And those are those videos that we've seen on social media, like this one,. Speaker C: When there's so many haters, I really don't care because their data has made me rich beyond my wildest dreams. Speaker B: Or this one that was of President Obama. Speaker D: We're entering an era in which our enemies can make it look like anyone is saying anything at any point in time, even if they would never say those things.

1:03 Speaker B: But actually, synthetic media is any media that's created by computers or modified by computers. Speaker B: So actually the media that we see around us, for example, playing around with Instagram filters or TikTok filters, all of those things are synthetic media because they're manipulated or they're creating or modified using the help of computers. Speaker B: So when you turn into a cat or you turn your mask on, that's actually synthetic media as well. Speaker B: And then there's also this high end synthetic media that the big movie producers use, you know, for making those masks, tracking people's faces and those very complex models in mapping those out and making them sound and look real in movies.

1:40 Speaker B: And there's a great development that's happening there. Speaker B: This is also synthetic media through this digital domain example here. Speaker B: So there's these two categories of the consumer side where, you know, we play around with the filters. Speaker B: You can use the Time Machine filter or Snapchat that I've used here when I turn myself older or younger, or they're those big high end movie productions where these professionally made used hours and months of work and millions of dollars to make happen. Speaker B: But now using those tools, those tools are getting more and more accessible.

2:08 Speaker B: We see the consumer ease of use coming to professional tools and professional tools becoming more and more available to consumers. Speaker B: And we see this middle persumer class coming up from companies like Synthesia, Rephrase or Windsor that we're going to see later on. Speaker B: And we're moving this general phase from digital to everything driven by AI. Speaker B: We already did that in music, where we moved from analog instruments to digital instruments. Speaker B: So now most of the music being produced, beats, voices, pitch corrections, all done by AI using computers.

2:39 Speaker B: Same thing happened in written press, from writing, sending fax to sending emails to now most of the emails being generated by computers using the titles and the content. Speaker B: And they're mostly automatically even sent out. Speaker B: And the same thing is happening now in the video world. Speaker B: We're no need for cameras, no need for microphones, lights or crews, and we can just generate all of that content. Speaker B: We can produce content on our computers. Speaker B: And I'm going to show you how I did the intro bit as well. Speaker B: But why now and why video?

3:05 Speaker B: First of all, why video? Speaker B: 68% Of consumers prefer watching videos instead of reading articles or looking infographics or ebooks. Speaker B: That's why YouTube is the second largest search engine. Speaker B: And the top three most popular apps for Gen Z are all video apps. Speaker B: Snapchat, Twitch and TikTok. Speaker B: And the consumer traffic by the end of this year will be 82% video. Speaker B: It's also important for business because it increases the click through rate, increases the exposure in Google search As they include 62% of videos in Google search results.

3:34 Speaker B: And people who have video on the website, they spend actually longer on the website, two minutes longer the next day. Speaker B: So here's an example. Speaker E: Everyone's got a personal video just for them. Speaker E: Hey Sebastian, how are you doing? Speaker E: Hey David, how are you doing? Speaker E: Hey Michael, how are you doing? Speaker E: And they just had to reply Using. Speaker B: Computer generated personalized email is actually having business impact as you saw before.

4:00 Speaker B: Peel, a company who's making phone cases used this as an example with 10,000 of their customers. Speaker B: The ones who got the personalized video from their CEO, thanking them from their purchase and the ones who didn't, and the ones who did get that personalized video, calling out their name and thanking them for the purchase were 87% more likely to buy from them again. Speaker B: So it's a significant business driver in getting more revenue and having a more personalized connection.

4:25 Speaker B: Synthesia themselves also reports that they're getting higher engagement when they're using video one of their clients and their massive cost savings in video production by using synthetic media production platforms. Speaker B: And what why now? Speaker B: Why we're just now getting those tools accessible to everybody? Speaker B: Number one thing is that Google machine learning for Word Accuracy surpassed human level accuracy in 2017, just four years ago. Speaker B: The same thing is now happening with.

4:50 Speaker A: The production of voices. Speaker B: So Google WaveNet is 92% human speech quality and being able to do that in 80 plus different languages and variants. Speaker B: So you can produce 20 seconds of audio in just one second of production. Speaker B: And this truly shows that we're passing the Turing test, which means that whether computer processes the audio or human processes the audio, we don't hear the difference. Speaker B: And I think we're hitting that inflection point now with some of the demos that we've seen from Google.

5:19 Speaker B: And also what else is happening is that image classification themselves. Speaker B: And the visual side is getting better. Speaker B: We passed human accuracy in 2015 and it is just getting better. Speaker B: We can use the machine learning to see people's emotions, recognize their faces and all the other attributes to then create engines to create people's faces. Speaker B: This is an Unreal Engine metahumans example to see how you can use their platform to generate all kinds of different faces. Speaker B: Or you can use platforms like thispersondoesonexist.com to generate as many fake faces as you want.

5:49 Speaker B: But what are the benefits? Speaker B: What is the reason of using these kinds of technologies? Speaker B: Number one, it's democratization. Speaker B: You're giving that ability to create videos to everybody. Speaker B: So you don't need a camera, it's accessible to everybody. Speaker B: It's much lower cost because you don't need to invest in that equipment. Speaker B: It's much more simpler. Speaker B: You don't know how to operate all of these platforms. Speaker B: So let's go into each one of them. Speaker B: How the content creation now drives scale, how it's much more personal, how it helps you reach more people by knowing or being able to share your content in more languages.

6:20 Speaker B: Number one is scale through automation. Speaker B: You know, if somebody purchases something on your website, you can then trigger it automatically through Zapier to go through, for example, someone like Synthesia or Windsor to trigger that personalized email to go to the customers to drive the customer love that we already saw for them coming back to you. Speaker B: The second thing is speed. Speaker B: You know, for me to record a 10 minute video takes 10 minutes. Speaker B: For AI to generate 10 minutes video,. Speaker A: Just 30 seconds, it's very, very fast.

6:45 Speaker B: It accelerates your production pipeline. Speaker B: And in order to create a digital version of my voice, for example, I descript in the first iteration, the first video that you saw, so I had to read 30 minutes of audio to the computer to then generate my voice. Speaker B: Custom voice model in Synthesia it cost thousand dollars to Create a custom avatar of yourself. Speaker B: And for that I needed to do five takes of three minute video and then I'll have an avatar that I can use indefinitely across all of my videos.

7:15 Speaker B: So how did it actually look like? Speaker B: Or how do those tools work? Speaker B: Because we talk about simplicity, right? Speaker B: This is an example of a descript where you can just type in the words or you can take existing script or existing video and then it automatically generates your voice, which you can then just export to be used across all of your distribution channels. Speaker B: Or you can use that voice recording to then put against your video. Speaker B: So here's what I'm going to do.

7:40 Speaker B: Within Synthesia and their interface, I'm showing how I created the intro video. Speaker B: First off, you go to their platform. Speaker B: There are so many options out there. Speaker B: It's like a true creation platform. Speaker B: So you can use their existing templates or you can import your own power PowerPoints to build your own templates. Speaker B: You can then also choose background, for example. Speaker B: In this case, I'm using the same background that I'm using for these videos to make it look very similar. Speaker B: You can size the avatar, you can add your text, you can add your graphics, you can add music, you can add different elements into the video and you can also change the person in the video.

8:11 Speaker B: Then in order to generate the voice, you can use their own voice generation platform, which is really good, actually even better than my voice model. Speaker B: And you can choose that in 55 different languages. Speaker B: In this case, I'm using the audio file that I generated already in descript as I want it to sound like me using my own voice, custom voice, and that's it. Speaker B: Once you hit generate, you just wait a couple of minutes and the video is going to be ready what you already saw in the intro bit. Speaker B: So we're moving from this investing 20 or hundreds of thousands of dollars to analog equipment to now being recently just able to couple of thousand to start being your filmmaking career.

8:45 Speaker B: By now just, you know, paying $30 in Synthesia platform and being able to create videos just using your computer with no need for cameras, lights or microphones. Speaker B: Here's an example of how it also makes it very personal. Speaker B: Messi, a famous football player. Speaker B: Here you can use your. Speaker B: You can just by entering some triggers like your name, your friend's name and where you want to see the game together with your friend, you can automatically generate videos that sound completely as it's coming from Messi and your friends won't notice the difference.

9:13 Speaker B: This is unbelievable, how good it is. Speaker F: Hey Stefan, what's up. Speaker F: My friend Sander has invited us to watch the game online. Speaker F: I hope I can make it if I can't be there. Speaker F: Enjoy the game. Speaker F: Ah, and don't forget to bring the snacks. Speaker F: Ciao. Speaker B: It's unreal how good it is, and I'm always amazed when I see examples like that. Speaker B: It can also help you reach a lot more people. Speaker B: In addition to being personal, you can also reach more people by having it available in many more languages.

9:41 Speaker B: Synthesia, for example, supports more than 50 languages. Speaker B: Here's an example of how it can be very powerful. Speaker B: In one of the campaigns that Synthesia did. Speaker B: Malaria isn't just any disease. Speaker B: It's the deadliest disease there's ever been. Speaker B: So imagine the reach that you can drive by being able to speak everybody's language around the world.

10:11 Speaker B: If you want to check out more examples and how all the other big companies that you see here are using their tools and go to Synthesia's website, which I've linked down below in this video. Speaker B: But what are the other use cases that synthetic media allows us to do, which we were not able to do before? Speaker B: A good example of that is the movement of VTubers, you know, where you can actually, rather than you being in the video, you can have a virtual character or your avatar being in the video. Speaker B: And here's example, one setup from Codemico, whose channel is also linked below.

10:41 Speaker C: My facial cam goes right here, okay. Speaker C: And it's an iPhone X. Speaker C: And so this is basically it. Speaker C: I have new fingers on. Speaker C: These are the gloves. Speaker C: See that? Speaker C: Thumbs up, whatever this is. Speaker C: Peace sign, 3, 4, 5. Speaker B: And this is incredible how much you can do with that.

11:07 Speaker B: You can create the whole virtual space, not just the character within that space. Speaker B: And of course, it's much more accessible. Speaker B: While this setup still cost like $10,000 for her, you can use Memojis and Animojis to animate yourself by just using your phone these days, by using the avatars. Speaker B: This has also sparked the start of virtual influencers. Speaker B: A company called Brute, which is in LA, has created Lil Miquella, who has got more than 3 million followers on Instagram.

11:34 Speaker B: By actually being completely virtual, it's also allowed AI companions to come around us. Speaker B: And it's not just, you know, being able to chat with them or talk in a chat environment and then reacting it on the screen, but also you can pick them out and actually place them within your space as a completely augmented reality experience where you can have a virtual friend, a companion that you can talk to anytime, that always Listens to you. Speaker B: It also sparked the start of digital humans by Unique, for example, where they create digital humans, but they also have created some that you can already interact with so you can go on their website and have a conversation with Einstein.

12:11 Speaker B: It's also used widely much more like closer example in learning and development in different companies, you know, to drive down the cost and make their content much more accessible and much more engaging. Speaker B: It's also used much in corporate communications, you know, where you just need to get a message across the whole company or things that move or change constantly. Speaker B: You can have those videos automatically generated or you can just use it, you know, for your own fun. Speaker B: We all know Reface app where you have the ton of templates where you can replace your face in any of the videos just to make them look, you know.

12:53 Speaker B: You know, those are just fun examples. Speaker B: But what are the risks by using those videos or those technologies across all of your video outputs and the use cases that we talked about? Speaker B: Number one certainly is ethics. Speaker B: While the tools are very powerful, it's important to keep people first always. Speaker B: And Synthesia is part of this content authenticity initiative. Speaker B: And I think whenever you're talking about these tools, you make sure that those companies have those ethics and principles in place.

13:19 Speaker B: You can only use the avatars with the person's permission, and they're not going to be shared publicly unless obviously people choose to do that. Speaker B: So ethics are key for those companies who have access to those technologies. Speaker B: But I think what is even more important than ethics for these companies is education and public knowledge. Speaker B: You know, we all got used to seeing costumes and then they're centuries old, from the sixth century before Christ, when we know that the person behind the costume is not the same person that they're acting out to be.

13:48 Speaker B: We're now used to that on TikTok and Instagram filters. Speaker B: You know, we know that's not real. Speaker B: We know that's actually generated by computer or lenses when somebody makes them look younger, older, or has different effects in them. Speaker B: You know, we got used to that in emails, you know, when 1970s when emails started, the first emails were sent, 99% of them were written by humans. Speaker B: While in 2000s, we're used to that receiving. Speaker B: The 99% of emails that we get are actually generated by computers using the name and the personalization and the titles and everything else that goes with an email.

14:18 Speaker B: So I think education is key. Speaker B: And while, you know, those cheaper tools are very accessible to everybody, they still don't look as good. Speaker B: Even the Pro Zoomer tools that don't look as good as something that we go and see in the movies because they take much longer time and much more effort. Speaker B: So there's really good defects that you see out there. Speaker B: Actually, somebody has to pay their time and energy to make those happen. Speaker B: There's usually somebody. Speaker B: Somebody's interest behind that.

14:43 Speaker B: But I think public knowledge is really the key. Speaker B: While public knowledge is key, I think it's going to be less and less easy to distinguish them, even with a higher awareness that those kinds of videos are out there. Speaker B: And I think fingerprinting is such a key area that needs to happen, whether on a device level or an application level where those videos are produced. Speaker B: And then with that, while we have the fingerprint, who's the original author of that video that's always attached and embedded into the file. Speaker B: We should have tools for consumers to then find out who is the owner.

15:10 Speaker B: You know, like in music, I hear a good song, I want to know who was the singer, who was the writer. Speaker B: Same way in YouTube, when you see a video that's using music, you could go out and validate, you know, who is the actual owner and who should claim the revenue from that video. Speaker B: So those tools already exist in music. Speaker B: They need to happen also for video. Speaker B: What does the future look like for video? Speaker B: I'm really excited about this part. Speaker B: We're moving from this idea that somebody looks like you, can act like you, but in the future can also think like you.

15:39 Speaker B: And I think this is a super powerful development. Speaker B: When somebody looks like you, we can make anybody say anything to use their likeness. Speaker B: And here's an example of digital domain Charlatan in the DHG group, we've been. Speaker G: Doing a lot of research, a lot of research on how to create digital humans, digital creatures, digital characters. Speaker B: There really is. Speaker G: There isn't any real way just to turn one person into somebody else. Speaker G: That technology just doesn't exist. Speaker G: We're not able to sort of take somebody's face and immediately just suddenly transform it into my face.

16:09 Speaker B: And there truly that technology does not exist. Speaker B: Never seen that. Speaker B: Of course it exists and it's actually live now. Speaker B: The second thing, to make things act like you. Speaker B: This gives us an ability to make things come alive. Speaker B: And a great example is the Google Lambda where they made paper airplanes and planet Pluto so you can converse with them. Speaker B: They have the knowledge of a Pluto and paper airplane from the world, from the web and so that you can now start having conversation with them.

16:35 Speaker B: Listen to a conversation the team had with Pluto a few days ago. Speaker H: I'm so curious about you, I sense your excitement. Speaker I: Ask me anything. Speaker H: Tell me what I would see if I visited. Speaker I: You would get to see a massive canyon, some frozen icebergs, geysers and some craters. Speaker H: It sounds beautiful. Speaker I: I assure you it is worth the trip.

17:00 Speaker I: However, you need to bring your coat because it gets really cold. Speaker H: I'll keep that in mind. Speaker H: Hey, I was wondering, have you ever had any visitors? Speaker B: And I think this is powerful. Speaker B: If you now put the likeness together with act like. Speaker B: So if you take the knowledge someone's likeness, how they look and how they sound, and then put that together with the knowledge, you get to autonomous digital humans. Speaker B: And this is another example of what Digital Domain is working at, where they created Duck as somebody you can converse with without necessarily knowing what they're going to say, because they're fully driven by their own AI in their looks, in their sound, in their knowledge.

17:35 Speaker B: Would you like to introduce yourself? Speaker J: Hello everyone. Speaker J: I'm an autonomous digital human digital replica of Doug Roble. Speaker J: Digital Domain has been on the forefront of visual effects for over two decades. Speaker J: These effects take thousands of hours and hundreds of skilled artists. Speaker B: But things have changed and things have truly changed. Speaker B: So imagine even in my case like creating a YouTube video without writing or filming.

18:02 Speaker B: Or you could just ask me, hey, can you tell me more about synthetic media? Speaker B: Then it automatically generates a video for you. Speaker B: So even today I can use a writer such as OpenAI GPT3 to write the script for sound. Speaker B: I can use Descript to then make it sound like me or someone like re speech or tool that was used in the recent Mandalorian series from Star Wars. Speaker B: And for camera and editing, you could use Synthesia, you know, a platform that brings it all together where you can add effects and text and then make it actually come out and sound really good.

18:32 Speaker B: So just maybe we are very soon. Speaker B: Of course it's going to take time getting to a place where you can start making Hollywood films with a laptop, which is the vision of for Synthesia company and the CEO as well. Speaker B: So thank you very much for watching. Speaker A: I hope you thank you from me as well. Speaker A: And just one more thing, please let me know down in the comments if you'd like me to create a video that is entirely not created by Sander, but AI. Speaker A: I'll ask GPT3 to write it, Descript to do the audio and Synthesia to film it.

19:02 Speaker A: Thanks and hope to see you next time.