Generative AI Explained: how your life & work will change

by Sander Saar

Speaker A: We've definitely hit the Apollo landing moment where all eyes are on artificial intelligence and how it will reset everything. S…. Along the way it covers Speaker A: The second thing that we talked ab…, Speaker A: They have Hundreds of billions of… and Speaker D: Just kidding, and more. Speaker A: I will then use, you know, Canva and Figma as a tool to create it. Speaker A: I will then put it on Behance, for example, as par….

For agents: machine-readable ARVP doc at /public/creators/sander-saar/videos/video_reLhJnQd-Lo/arvp.json · Browse this video library

Key findings and exact moments

  1. 0:00Speaker A: We've definitely hit the Apollo la…

  2. 4:32.79Speaker A: The second thing that we talked ab…

  3. 9:05.82Speaker A: They have Hundreds of billions of…

  4. 12:39.02Speaker D: Just kidding

  5. 16:24.42Speaker A: Audio is another big area of model…

  6. 20:08.92Speaker A: So let's see some of the examples…

  7. 24:25.42Speaker A: So for example, taking a image of…

  8. 29:31.4Speaker A: I will then use, you know, Canva a…

Chapters

  1. 0:00 Speaker A: We've definitely hit the Apollo la…

    Speaker A: We've definitely hit the Apollo landing moment where all eyes are on artificial intelligence and how it will reset everything. S…

    On screen: A man speaks to the camera beside an evolving diagram illustrating the hierarchical layers of generative AI interfaces.

  2. 4:32 Speaker A: The second thing that we talked ab…

    Speaker A: The second thing that we talked about, why it's happening in addition to interface is cost. Speaker A: We talked about in my pre…

    On screen: The presenter stands against an animated graphic comparing generative AI tools and the cost shifts across media and education.

  3. 9:05 Speaker A: They have Hundreds of billions of…

    Speaker A: They have Hundreds of billions of parameters. Speaker A: Baidu has one of the largest models out there with 260 billion paramete…

    On screen: A presenter speaks in front of a diagram illustrating various practical applications of language models.

  4. 12:39 Speaker D: Just kidding

    Speaker D: Just kidding. Speaker D: That's. Speaker D: That's not really a thing. Speaker A: And that's something where it becomes really p…

    On screen: A man speaks facing the camera as animated informational columns detail the data, training, model, and inference pipeline of Stable Diffusion on a white background.

  5. 16:24 Speaker A: Audio is another big area of model…

    Speaker A: Audio is another big area of modeling, specifically in speech modeling and music, where Google has been building large language…

    On screen: The host speaks to the camera in front of a chart classifying various generative AI audio use cases and tools.

  6. 20:08 Speaker A: So let's see some of the examples…

    Speaker A: So let's see some of the examples and the first one here is the Runway ML and their new model to stylize any content, generate v…

    On screen: A presenter gestures while speaking in front of a digital slide displaying various scientific and healthcare applications of artificial intelligence.

  7. 24:25 Speaker A: So for example, taking a image of…

    Speaker A: So for example, taking a image of your retinal scanned and then matching it to a diabetics database to be much more accessible a…

    On screen: The presenter stands in front of a white background with comparison bar charts showing metrics for Google, ChatGPT, and Stable Diffusion.

  8. 29:31 Speaker A: I will then use, you know, Canva a…

    Speaker A: I will then use, you know, Canva and Figma as a tool to create it. Speaker A: I will then put it on Behance, for example, as par…

    On screen: A man speaks facing the camera with a graphic on screen detailing creative workflow tools before and after AI.

Visual moment index

What the footage shows, shot by shot. It is derived from the video itself, independent of the commentary. Timestamps link to the exact second.

0:00–0:09 A host and his AI avatar clone appear side by side in a studio setting in front of an image of an astronaut and moon rover. On screen: AI WILL RESET EVERYTHING

0:09–0:18 A man enthusiastically speaks and gestures in front of a slide depicting an astronaut and rover on the moon with bold text. On screen: AI WILL RESET EVERYTHING

0:18–0:40 A presenter speaks in the lower-left corner beside an infographic illustrating a queue of job categories representing AI impact perceptions from five years ago. On screen: AI IMPACT 5 YEARS AGO 1. PHYSICAL LABOR 2. LESS DEMANDING COGNITIVE LABOR 3. HIGHLY DEMANDING COGNITIVE LABOR 4. CREATIVITY STILL FUMS

0:40–0:42 The presenter talks in the corner next to a graphic depicting a queue of four professions facing AI disruption. On screen: AI IMPACT NOW 1. CREATIVITY 2. HIGHLY DEMANDING COGNITIVE LABOR 3. LESS DEMANDING COGNITIVE LABOR 4. PHYSICAL

0:45–0:50 The presenter continues explaining AI impact tiers alongside the queue illustration of diverse workers. On screen: AI IMPACT NOW 1. CREATIVITY 2. HIGHLY DEMANDING COGNITIVE LABOR 3. LESS DEMANDING COGNITIVE LABOR 4. PHYSICAL

0:50–1:09 A man addresses the camera beside a comparative bar chart illustrating the time taken by various platforms to reach 100 million users. On screen: TIME TO REACH 100 MILLION USERS CHATGPT 2 MONTHS TIKTOK 9 MONTHS FACEBOOK 4 YEARS MOBILE 12 YEARS PC 14 YEARS POKEMON GO <1 MONTHS

1:09–1:13 A presenter speaks in the foreground in front of a Google Trends chart displaying interest in artificial intelligence over time. On screen: Google Trends Explore Artificial intelligence Topic + Compare Worldwide Past 5 years All categories Web Search Interest over time 25 50 75 100 25 Feb 20... 13…

1:13–1:21 The presenter continues talking against a graphic timeline illustrating the PC era and early browser search. On screen: PC 1980 BROWSER SEARCH

1:21–1:32 A man in a patterned shirt speaks beside a timeline graphic presenting the PC and Mobile computing eras. On screen: MOBILE 2007 PC 1980 SOCIAL & SUPER APPS APP STORES BROWSER SEARCH

1:32–1:51 A man speaks to the camera as an infographic beside him expands to showcase the evolution from PC to AI eras. On screen: CLOUD 2010 CROSS-PLATFORM COLLABORATIVE MOBILE 2007 SOCIAL & SUPER APPS APP STORES PC 1980 BROWSER SEARCH AI 2015 INTELLIGENT AGENTS ?

1:51–1:58 The presenter is displayed on the bottom-left alongside an embedded video of Sam Altman speaking into a microphone. On screen: WHY NOW? Felic

1:58–2:30 Sam Altman speaks into a microphone while seated on a stage during an interview. On screen: SAM ALTMAN STRICTLY VC Felic

2:30–2:37 A presenter gestures enthusiastically in front of a graphic diagram illustrating the key drivers of the AI paradigm shift. On screen: KEY DRIVERS OF AI PARADIGM SHIFT QUALITY INTERFACE COST

2:37–2:47 A presenter sits in the bottom left corner talking energetically alongside a graphic slide showing a Venn diagram of AI paradigm shift drivers. On screen: KEY DRIVERS OF AI PARADIGM SHIFT QUALITY INTERFACE COST

2:47–3:06 A man speaks towards the camera alongside a presentation comparing the GPT-3 Playground interface to ChatGPT. On screen: INTERFACE GPT-3 PLAYGROUND CHATGPT HUMAN LEVEL CHATBOT INTERFACE JUL 2020 NOV 2022

3:06–3:20 The presenter stands in front of a white presentation slide displaying GPT-3, OpenAI branding, and a base model diagram. On screen: INTERFACE GPT-3 BASE MODEL OpenAI

3:20–3:32 A presenter speaks in front of an educational graphic illustrating AI layers for InstructGPT and GPT-3. On screen: INTERFACE InstructGPT GPT-3 INSTRUCT BASE MODEL OpenAI

3:32–4:17 A man speaks to the camera beside an evolving diagram illustrating the hierarchical layers of generative AI interfaces. On screen: INTERFACE ChatGPT InstructGPT GPT-3 EXCHANGE INSTRUCT BASE MODEL OpenAI Bing RETRIEVAL ... ACTION

4:17–4:33 A presenter speaks in the foreground against a slide presentation displaying Replit Ghostwriter Chat and Bing in Edge interface examples. On screen: INTERFACE REPLIT GHOSTWRITER CHAT BING IN EDGE Conversational AI in your IDE Trying asking questions about your code or follow-ups about a previous message.…

4:33–4:52 The presenter stands on the left of the frame gesturing while a comparison graphic of video production costs across Analog, Digital, and AI formats is displayed. On screen: COST COST OF MAKING VIDEOS Analog Digital AI $200,000 $2,000 $20

4:52–4:53 A man in a patterned shirt speaks to the camera against a digital backdrop showing a person editing video on a dual-monitor laptop setup. On screen: COST Make Hollywood films with a laptop

4:53–5:27 The presenter stands against an animated graphic comparing generative AI tools and the cost shifts across media and education. On screen: COST Writer OpenAI GPT-3 Sound descript respeeCHer Camera & Editing Synthesia COST OF INTELLIGENCE College $250,000 Yale Cornell BROWN PRINCETON Online $15,000…

5:27–5:47 A speaker presents in the corner beside a graph depicting decreasing compute cost and increasing data volume over time for AI models. On screen: COST Volume Time DATA GPT-3 (175B) Megatron-LM (8.3B) Megatron-Turing NLG (530B) Turing-NLG (17.2B) T5 (11B) GPT-2 (1.5B) BERT-Large (340M) ELMo (94M) Amazon…

5:47–5:53 The presenter gestures enthusiastically beside an infographic comparing physical atoms in a warehouse to digital bits on a computer screen. On screen: COST ATOMS VS BITS

5:53–6:02 The host speaks in the foreground beside a graphic comparing 'ATOMS' in a warehouse against 'BITS' on a computer screen. On screen: COST ATOMS vS BITS

6:02–6:24 A man speaks on the bottom left corner next to an informational graphic explaining algorithmic models with input and output flowcharts. On screen: QUALITY 1990-2000 ALGORITHMIC MODEL HUMAN ITIZATION DIGITIZATION Input Set of rules of\nobtain the expected\noutput from the\ngiven input. Output Algorithm

6:24–6:41 A presenter speaks in the lower left against a white presentation slide detailing the 2000-2010 era of data science and big data architecture. On screen: QUALITY 1990-2000 ALGORITHMIC MODEL HUMAN DIGITIZATION 2000-2010 DATA SCIENCE BIG DATA Source Data Data Replication & Ingestion Data Lake & Data Warehouse…

6:41–7:10 A speaker appears in the foreground beside an infographic outlining the evolution of machine learning and image classification accuracy. On screen: QUALITY 1990-2000 ALGORITHMIC MODEL HUMAN GITIZATION 2000-2010 DATA SCIENCE BIG DATA 2010-2020 MACHINE LEARNING HUMAN + COMPUTER DEEP Image Classification…

7:10–7:35 A man speaks and gestures in front of an infographic illustrating the evolution of AI models across decades. On screen: QUALITY 1990-2000 ALGORTHMIC MODEL HUMAN IZATION 2000-2010 DATA SCIENCE BIG DATA 2010-2020 MACHINE LEARNING HUMAN + COMPUTER NEURAL NETWORKS 2020-... AI MODELS…

7:35–7:53 A presenter speaks energetically alongside an infographic slide displaying a Venn diagram of key drivers of the AI paradigm shift with ChatGPT at the center. On screen: KEY DRIVERS OF AI PARADIGM SHIFT HUMAN LEVEL INTELLIGENCE ChatGPT CHAT FREE

7:53–8:01 A man in a patterned shirt speaks in the bottom left against a presentation slide showing a portrait labeled 'PASSIVE'. On screen: PASSIVE

8:01–8:19 A presenter speaks in the lower left while a slide diagram illustrating stages of human labor—from Passive to Narrator—progressively populates. On screen: PASSIVE LABORER CREATOR NARRATOR

8:19–8:25 A male speaker gestures while discussing an on-screen chart depicting stages of human productivity from passive to narrator. On screen: PASSIVE LABORER CREATOR NARRATOR

8:25–8:39 A presenter speaks in front of a large digital graphic showcasing six domains of generative AI: language, images, audio, video, code, and science. On screen: LANGUAGE IMAGES AUDIO VIDEO CODE SCIENCE

8:39–9:17 A presenter speaks against a backdrop displaying a chart of major large language models and their developer companies. On screen: LANGUAGE 175B ChatGPT & GPT-3 OpenAI Bing Chat Microsoft 137B Bard Google 52B Claude ANTHROP\C 260B ernie Bot Baidu 27B Carper AI Alibaba.com aws 70B…

9:17–9:37 A presenter sits on the lower left of the frame gesturing while explaining data sources next to an infographic graphic labeled 'DATA' with various platform icons and an OpenAI logo. On screen: LANGUAGE OpenAI DATA WIKIPEDIA ommon Crawl 45 TB

9:37–10:01 The speaker appears on the bottom left corner presenting a diagrammatic pipeline of OpenAI's language model workflow, including data collection, training, model size, and inference. On screen: LANGUAGE OpenAI DATA Wikipedia mmon Crawl 45 TB TRAINING Azure $4.6M MODEL GPT-3 800 GB INFERENCE $0.02-8/result

10:01–10:10 A man in a patterned shirt speaks next to a graphic detailing language model data, training, model size, and inference costs. On screen: LANGUAGE DATA WIKIPEDIA ommon Crawl 45 TB TRAINING Azure Text Prediction Task Classifier Layer Norm Feed Forward Masked Multi Self Attention Text & Position…

10:10–10:25 A man gestures expressively on the left side of the screen against a background displaying a spreadsheet of AI models from LifeArchitect.ai. On screen: LANGUAGE 2023 LifeArchitect.ai data (shared) File Edit View Insert Format Data Tools Extensions Help Comment only Share A1 fx Model Lab Selected playgrounds…

10:25–11:27 A presenter speaks in front of a diagram illustrating various practical applications of language models. On screen: LANGUAGE Text Generation Gap Inc. Summarization Translation Welcome to the new Bing Search Duplicate bot0 as bot2 and move the copy to its left Intelligent…

11:27–11:34 The host speaks in the foreground against a background screen showing Jasper AI's writing templates dashboard. On screen: LANGUAGE Dashboard Templates Recipes Community Documents Trash AI outputs Favorites Search templates All Frameworks Email Website Blog Ads Ecommerce Social…

11:34–11:37 The host continues speaking while the background displays an animated 3D robot mascot. On screen: LANGUAGE

11:37–11:38 An over-the-shoulder view shows a computer screen presenting Jasper AI's writing interface. On screen: Blog Meet Jasper, The Future of Writing Blog Articles.

11:38–11:40 A woman looks intently at her computer monitor inside a brightly lit office. On screen: JASPER JASPER.AI

11:41–11:44 A digital screen displays Jasper's dashboard and AI content generation tools. On screen: Let Jasper help you write longer articles from start to finish. Tell Jasper exactly what to write with a command Updated 4h ago Paragraph Generator Generate…

11:44–11:45 A 3D animated white and blue companion robot stands against a blurred office background. On screen: JASPER JASPER.AI

11:45–11:48 A woman sits at her office desk looking attentively next to a small desktop robot figurine. On screen: JASPER JASPER AI

11:49–11:51 A small white and blue robot gestures with its hands while speaking on a table. On screen: JASPERPER.AI

11:51–11:53 Text is being generated automatically on a white digital document screen. On screen: JASPER and that's why the perfect pair of socks is like a hug for Jasper

11:53–11:54 A woman smiles warmly at her desk while looking at a red computer monitor next to a small robot figurine.

11:54–12:17 The presenter is framed in the lower-left corner while presenting a video clip of Ryan Reynolds reading a Mint Mobile script generated by AI. On screen: mintmobile JOKE CURSE WORD STILL GOING BIG WIRELESS

12:17–12:21 Ryan Reynolds looks at a sheet of paper while seated in front of a wood-paneled wall with text overlays. On screen: mintmobile JOKE CURSE WORD STILL GOING BIG WIRELESS RYAN REYNOLDS Mint Mobile Owner & User MINT MOBILE

12:21–12:31 Ryan Reynolds reads lines from a script as green checkmarks appear beside checklist points on screen. On screen: mintmobile RYAN RENOLDS | MINT MOBILE JOKE CURSE WORD STILL GOING BIG WIRELESS

12:31–12:36 A mint green promotional screen displaying deal terms and conditions. On screen: mintmobile RYAN RENOLDS | MINT MOBILE JOKE CURSE WORD STILL GOING BIG WIRELESS BUY 3 MONTHS, GET 3 MONTHS FREE Offer expires 1/15/23. New customers only.…

12:36–12:42 Ryan Reynolds wearing glasses reads from a sheet of paper against a rustic wood wall. On screen: mintmobile RYAN RENOLDS MINT MOBILE JOKE CURSE WORD STILL GOING BIG WIRELESS

12:42–12:50 A man gestures and talks enthusiastically in front of a background graphic displaying the Microsoft Bing search homepage. On screen: Microsoft Bing Chat Work Images Videos Shopping Maps LANGUAGE Takumi Morita 8,750 Ask me anything... 0/1000 SMASH

12:50–13:27 A presenter in the lower-left corner speaks beside an animated display demonstrating the Bing Chat interface and search features. On screen: Arts and crafts ideas, with instructions for a toddler using only cardboard boxes, plastic bottles, paper and string Ask me anything Give me a workout I can do…

13:27–13:51 The host speaks to the camera while an infographic displaying various AI image generation models and their creators is featured in the background. On screen: IMAGES 3.5B DALL-E 2 OpenAI 20B Parti & Imagen Google aws 890M Stable Diffusion stability.ai Midjourney

13:51–14:37 A man speaks facing the camera as animated informational columns detail the data, training, model, and inference pipeline of Stable Diffusion on a white background. On screen: IMAGES stable diffusion DATA LAION-5B gettyimages 5B TEXT-IMAGE PAIRS 120 TB TRAINING $600K MODEL 2 GB

14:41–15:19 A man speaks to the camera against a grey background alongside infographics demonstrating image generation capabilities and applications. On screen: IMAGES stable diffusion DATA TRAINING MODEL INFERENCE LAION-5B TEXT-IMAGE PAIRS 120 TB $600K 2 GB Draw Things Diffusion Bee FREE Text-to-image Image-to-image…

15:19–15:23 The presenter stands in front of a digital slide displaying interconnected image nodes while introducing examples. On screen: IMAGES

15:23–15:27 A black screen displays the DALL-E 2 OpenAI logo followed by the text prompt 'a koala dunking a basketball'. On screen: DALL-E 2 OPENAI a koala dunking a basketball

15:27–15:32 A visual demonstration transitions from static noise into a rendered image of a koala dunking a basketball. On screen: DALL-E 2 OPENAI

15:33–15:39 A digital screen demonstrates DALL-E 2's inpainting feature by erasing a dog on a chair and typing a prompt for a cute cat. On screen: DALL-E 2 OPENAI cut cute cat

15:39–15:45 The masked image generates a realistic cat resting on the armchair to demonstrate AI image completion. On screen: DALL-E 2 | OPENAI cute cat

15:45–15:50 The presenter stands in the lower left corner against a backdrop displaying mobile UI mockups generated by AI. On screen: IMAGES DALL-E 2 OPENAI 9:41 Settings John Smith Your display name +1 (123) 456-7890 Your phone number Change Password Delete account Dan Brown Bestselling…

15:50–15:56 The presenter talks while a prompt input bar types a request for a dog walking app UI against a dark background. On screen: Describe your design... Generate An onboarding screen for a dog walking a GALILEO AI USEGALILEO.AI An onboarding screen for a dog walking app

15:56–16:05 The host speaks in the bottom left corner as a demo of Galileo AI generating mobile app UI screens is displayed. On screen: GALILEO AI | USEGALILEO.AI 9:41 Welcome to Doggy Daycare We'll take care of your pup while you're away. Let's get started! Let's Go Describe your design...…

16:05–16:12 A demonstration of Galileo AI generating mobile user interface screens is displayed alongside the speaker in the lower-left corner. On screen: GALILEO AI USEGALILEO.AI 9:41 Welcome to Doggy Daycare We'll take care of your pup while you're away. Let's get started! Let's Go Settings John Smith Your…

16:15–16:17 The speaker speaks next to text on a black background reading that designs generated with AI are editable in Figma. On screen: GALILEO AI USEGALILEO.AI 9:41 Doggy p while you're Settings John Smith Your display name +1 (123) 456-7890 Your phone number Change Password Delete account Dan…

16:17–16:20 A slide categorizing generative audio AI models across speech and music appears next to the speaker.

16:20–16:55 The presenter stands in the lower left against a white graphic diagramming various audio AI models for speech and music generation. On screen: AUDIO Speech AudioLM WaveNet Google aws Polly Whisper OpenAI Speechify Music Music LM SingSong MuseNet UberDuck UBERDUCK

16:55–17:35 The host speaks to the camera in front of a chart classifying various generative AI audio use cases and tools. On screen: AUDIO Ultra-realistic voice cloning with Overdub Text-to-speech OpenAI Whisper Robots with Human-Level Speech Recognition Skills? Speech-to-text Voice Cloning…

17:35–17:48 A presenter speaks in the lower left corner against a grey background next to a graphic showing a podcast episode interface from Podcast.ai. On screen: AUDIO PODCAST.AI - EPISODE 1 Joe Rogan interviews Steve Jobs 00:04|19:17 SUBSCRIBE SHARE MORE INFO Transistor 4 LATEST EPISODES Zach Galifianakis talks movies…

17:48–18:07 A dramatic black-and-white close-up portrait of Steve Jobs resting his chin in his hand against a black background. On screen: PODCAST.AI PLAY.HT

18:07–18:28 The presenter speaks animatedly in front of a split screen featuring a concert clip and logos for ChatGPT and Uberduck. On screen: AUDIO CHAT GPT UBERDUCK

18:28–18:40 David Guetta performs behind a DJ mixer on stage, gesturing to an energized crowd illuminated by green stage lights. On screen: DAVID GUETTA UBERDUCK Rollin' In My 64 DAVID GUETTERDUCK

18:40–18:43 A vertical concert video captures a massive crowd in an arena bathed in green strobe lighting from behind the DJ booth. On screen: DAVID GUETTA UBERDUCK

18:43–18:45 David Guetta speaks enthusiastically while gesturing with his hands during an interview. On screen: DAVID GUETTA UBERDUCK Eminem, bro!

18:45–18:46 David Guetta smiles and continues explaining a story during the seated interview. On screen: DAVID GUETTA UBERDUCK There’s something that I made

18:46–18:50 A framed vertical video clip shows music producer David Guetta talking and smiling during an interview. On screen: DAVID GUETTA UBERDUCK And it worked so good!

18:50–18:54 David Guetta continues speaking in the vertical interview footage about discovering AI websites. On screen: DAVID GUETTA UBERDUCK I discovered those websites,

18:54–19:32 The presenter stands in front of a chart categorizing generative AI video tools into 3D and avatar generation. On screen: VIDEO 3D Imagen Video Google aws Make-A-Video Meta Krikey AI krikey Runw ru Avatars D-ID Synthesia synthesia Rephrase Rephrase.ai

19:32–20:08 A man speaks enthusiastically while standing in front of a presentation slide displaying various generative AI video use cases. On screen: VIDEO Text-to-video Image-to-video Personalized video Text-to-3D Classification Enhancements & Edit Make-A-Video with text Robot dancing in times square.…

20:08–20:14 The host speaks in front of a video demonstration showcasing a street scene with a text prompt interface. On screen: VIDEO Make it look more cinematic Add

20:14–20:17 A video clip displays a man filming a selfie video outdoors as an AI prompt overlay types 'Claymation style'. On screen: Claymation s Generate

20:17–20:20 The video transforms the previous street footage into a stylized claymation animation. On screen: GEN-1 RUNWAYML

20:22–20:23 A video clip displays an astronaut on the moon inset over a man walking, illustrating a visual prompt for Gen-1. On screen: GEN-1 RUNWAYML

20:23–20:26 A demonstration video shows a person walking turned into an astronaut walking across a rocky moon landscape. On screen: GEN-1 RUNWAYML

20:26–20:32 A woman looking forward is transformed into an animated metallic bearded sculpture using Gen-1 video-to-video AI. On screen: GEN-1 | RUNWAYML

20:34–20:39 A male presenter in the lower left corner talks against a screen recording of the D-ID web interface. On screen: D-ID VIDEO A portrait of 19th century portrait of a happy victorian gentleman CREATE Type Your Text Upload Voice Audio 514 characters left English - US Amber…

20:39–20:46 A motion graphics screen capture demonstrates the D-ID Creative Reality Studio software creating animated avatars. On screen: D-ID A portrait of | CREATE Type Your Text Upload Voice Audio 514 characters left Language English - US Voice Amber (Neural) Voice Style Cheerful Choose a…

20:46–20:51 Motion graphics and screencast display D-ID's interface showcasing AI avatars, Stable Diffusion, and GPT-3 integration. On screen: D-ID REALITY STUDIO using Stable Diffusion GPT-3

20:51–21:03 A screencast demonstration shows the D-ID Reality Studio interface generating three custom AI portrait avatars from descriptive text prompts. On screen: D-ID REALITY STUDIO A portrait of | photo r CREATE Type Your Text Upload Voice Audio 514 characters left Language English - US Voice Amber (Neural) Voice Style…

21:03–21:15 A man speaks against a background displaying the Krikey AI Animated Avatars website interface. On screen: Krikey.ai Intro Product Pricing Documentation Use Cases About Us FAQs Get Started VIDEO Due to high demand, you may experience some waiting time while…

21:15–21:17 A screen capture showing text being typed into an AI prompt box on a dark interface. On screen: Describe what you want to see Cool dance moves Generate

21:17–21:18 A 3D animated character dances on screen alongside promotional graphic text. On screen: ANIMATED AVATARS KRIKEY AI ONE TOOLKIT TO REC

21:18–21:21 Animated kinetic typography text flashes across a dark digital tunnel background with circular glowing light trails. On screen: ANIMATED AVATARS KRIKEY AI BUILD CUSTOMIZE LAUNCH

21:21–21:23 The Krikey AI logo icon and brand name are displayed against a dark geometric grid background. On screen: ANIMATED AVATARS KRIKEY AI

21:23–21:26 A promotional graphic presents Krikey AI SDK offerings surrounded by illustrated feature cards on a blue gradient backdrop. On screen: Casual Gaming SDK ANIMATED AVATARS KRIKEY AI Dynamic 3D NFT Minting Character Builder SDK SDKs we offer Describe what you want to see Generate AI Assets…

21:26–21:29 Futuristic neon light streaks burst outward with dynamic 3D text overlays advertising Krikey AI animation tools. On screen: ANIMATED AVATARS KRIKEY AI MAKE GAMES, FILM AVATAR ANIMATIONS, METAVERSE CONTENT

21:29–21:31 Graphic display showing the user interface of an AI text-to-animation tool generating 3D avatar movements from prompts. On screen: AI Text to Animation Tool ANIMAT 3D Assets Use AI to generate 3D Assets on or off chain Describe what you want to see E.g Dog Dressed for Indian wedding…

21:31–21:35 The host speaks to the camera in the foreground beside an inset video clip showcasing Generative AI on Roblox. On screen: VIDEO Generative AI on ROBLOX

21:35–21:37 The presenter is inset in the bottom left corner next to a slide titled 'Generative AI on ROBLOX' showing 3D material sphere assets. On screen: Generative AI on ROBLOX

21:38–21:43 A demonstration of Roblox AI Materials changing car surface textures via text prompts, alongside the inset speaker. On screen: AI Materials brushed metal diamond plate pattern| Generate purple foil, crumpled pattern, reflective| ROBLOX GENERATIVE AI red paint, reflective glossy finish|

21:43–21:49 A demonstration showing text-guided scripting in Roblox turning on a red sports car's headlights. On screen: lights on! ROBLOX | GENERATIVE AI -- Toggle the headlights when the user presses 'H' -- Blink the headlights when the user pres…

21:49–21:52 A red sports car drives down a dark, wet city street in a video game demo with the presenter inset in the lower-left corner. On screen: BrokenToy09 50 MPH ROBLOX GENERATIVE AI

21:52–21:59 Generated code is superimposed on the screen next to a red car following a text prompt to create rain. On screen: ROBLOX GENERATIVE AI -- Add an exp make it rain -- Add an exponential rain tween with a time of 1 -- Add a linear splash tween with a time of 3 -- A local…

21:59–22:02 The red car begins floating in mid-air in response to a text command prompt, while the host talks in the lower-left PIP. On screen: make it float AI GEN material ROBLOX GENERATIVE AI

22:03–22:04 The speaker appears in a small lower-left overlay while the main display shows the Roblox logo animation on a dark background. On screen: ROBLOX

22:04–22:28 A man speaks in the lower left against a slide showing AI coding tools from OpenAI, GitHub, Replit, and Google. On screen: CODE Codex OpenAI CoPilot GitHub Ghostwriter replit ... Google

22:28–22:47 Speaker A presents in the corner of a slide illustrating four primary coding use cases: autofill, text-to-code, QA & reviews, and code-to-text. On screen: CODE Autofill Text-to-code QA & Reviews Code-to-text

22:47–22:50 A presenter speaks in the lower-left overlay before a demonstration of Replit's code editor and AI debugger interface. On screen: GwChat amanm3 CODE Stop Invite index.js const express = require('express'); const app = express(); const path = require('path'); app.get('/', (req, res) => {…

22:50–22:54 A close-up view of the Ghostwriter Replit chat screen with the presenter speaking in the bottom left corner. On screen: GwChat amanm3 Run GHOSTWRITER CHAT | REPLIT index.js app.listen() callback const express = require('express'); const app = express(); app.get('/', (req, res)…

22:54–22:56 The screen switches back to the full Replit IDE view showing Ghostwriter answering a question, with the presenter talking at the bottom left. On screen: GwChat amanm3 Run GHOSTWRITER CHAT REPL index.js app.listen() callback const express = require('express'); const app = express(); app.get('/', (req, res) => {…

22:56–22:58 A split-screen software interface displays code editing and a console alongside the presenter in the lower left corner. On screen: sInit (/home/runner/GwChat/node_modules/express/li init.js:40:5) r: path is not defined runner/GwChat/index.js:6:16 handle [as handle_request]…

22:58–23:03 A close-up view focuses on the Ghostwriter console displaying code debugging suggestions while the speaker explains the interface. On screen: GHOSTWRITER CHAT REPLIT at Function.process_params (/home/runner/GwChat/node_modules/express/lib/router/index.js: 346:12) at next 280:10) at expressInit…

23:03–23:11 The full code editor, web preview, and chat window are displayed as the user interacts with the Ghostwriter assistant. On screen: GwChat amanm3 Stop GHOSTWRITER CHAT REPLIT index.js app.listen() callback const express = require('express'); const app = express(); app.get('/', (req, res) =>…

23:11–23:15 The speaker appears in a bottom-left overlay explaining code generation while a Replit Ghostwriter Chat screen recording is shown. On screen: GwChat amanm3 GHOSTWRITER CHAT REPLIT Stop index.js 1 2 3 4 5 6 7 8 9 10 11 12 13 14 const express = require('express'); const app = express(); const path =…

23:15–23:24 The speaker talks in a corner overlay in front of a GitHub Copilot promotional slide displaying coding productivity metrics. On screen: CODE AI pair programming is here. 75% more fulfilled 55% faster coding runtime.go JS days,between 9 // Get average runt 10 func averageRuntime 11 var totalTime…

23:24–23:49 A presenter speaks against a grey backdrop displaying a slide on AI science applications including AlphaFold 2, ESM Fold, RoseTTAFold, and LibreFold. On screen: SCIENCE AlphaFold 2 DeepMind aws ESM Fold Meta RoseTTAFold NIH National Institutes of Health LibreFold OpenBioML

23:52–24:50 A presenter gestures while speaking in front of a digital slide displaying various scientific and healthcare applications of artificial intelligence. On screen: SCIENCE Review & Analysis Predictive modelling Precision Medicine NAD Search NAD-specific glutamate dehydrogenase MYCOBACTERIUM TUBERCULOSIS (STRAIN ATCC 25618…

24:50–24:52 A blue screen displays text reading 'Protein folding problem' alongside animated diagrams and an incoming cursor. On screen: Protein folding problem

24:52–24:54 A text selection cursor highlights and replaces the headline text with 'AlphaFold' against a blue background. On screen: ALPHAFOLD DEEPMIND Protein folding problem

24:54–24:56 Motion graphics zoom through speed lines to reveal a white screen with an animated trajectory arrow. On screen: ALPHAFOLD DEEPMIND

24:56–25:04 Animated lines and geometric shapes grow and fold dynamically across a clean white background labeled with the AlphaFold and DeepMind logo. On screen: ALPHAFOLD DEEPMIND

25:07–25:22 A minimalist animation transitions from protein molecules in a dish to an illustrative cityscape of science and medicine, followed by recycling plastic bottles floating in water. On screen: ALPHAFOLD DEEPMIND

25:22–25:25 An animated graphic displays a circular arrow symbol emerging over an orange dotted cloud shape alongside the AlphaFold DeepMind logo. On screen: A

25:25–25:40 The presenter stands in front of a slide displaying six key challenges for AI adoption. On screen: CHALLENGES FOR AI ADOPTION HALLUCINATIONS & OUTPUT ERRORS HIGH COST LACK OF LONG TERM MEMORY SLOW INFERENCE TIME DATA ACCESS & COPYRIGHT CLOSED SYSTEMS & BIASES

25:40–25:42 A man speaks at the bottom left against a backdrop collage of news headlines about AI chatbots. On screen: HALLUCINATIONS Intelligence An Unsettling Chat With Bing Read the Conversation How Chatbots Work Spotting A.I.-Gen The New York Times THE SHIFT A Conversation…

25:42–25:52 The presenter continues talking in front of the AI hallucination news article collage. On screen: HALLUCINATIONS The New York Times An Unsettling Chat With Bing Read the Conversation How Chatbots Work Spotting A.I.-Generated ARTIFICIAL INTELLIGENCE / TECH /…

25:52–25:56 A presenter speaks beside an inset video showing a chimpanzee aggressively interacting with a mirror in a forest. On screen: HALLUCINATIONS

25:56–26:04 The presenter is keyed into the lower-left corner over a full-screen video of a chimpanzee violently thrashing branches at its reflection in a mirror outdoors. On screen: Prof. Anderson HUBERT BRIERRE

26:05–26:12 A video clip of a chimpanzee discovering a mirror in a lush forest is shown with the speaker talking in an inset window. On screen: Prof. Anderson Xavier HUBERT BRIERRE the calls of the chimpanzees in the backgro

26:12–26:18 A chimpanzee charges and attacks its own reflection in a freestanding outdoor mirror in the jungle. On screen: Prof. Anderson Xavier HUBERT-BRIERRE

26:18–26:24 A presenter speaks in an inset frame while a chimpanzee reacts aggressively to its own reflection in a jungle mirror experiment in the main video. On screen: Prof. Anderson Xavier HUBERT BRIERRE

26:24–26:39 The presenter speaks to the camera in the foreground against a background screenshot of the Bing AI chat interface. On screen: Microsoft Bing SEARCH CHAT HALLUCINATIONS Welcome to the new Bing Your AI-powered answer engine Ask complex questions Get better answers Get creative…

26:39–27:11 The presenter gestures and speaks while framed next to an on-screen graphic comparing output errors in Google Bard and OpenAI ChatGPT/Bing. On screen: OUTPUT ERRORS Google what new discoveries from the James Webb Space Telescope can I tell my 9 year old about? Your 9-year old might like these recent…

27:11–27:23 A presenter speaks while gesturing next to a graphic showing the logos for ChatGPT and WolframAlpha. On screen: OUTPUT ERRORS ChatGPT WolframAlpha

27:23–27:26 The presenter explains AI integrations in the lower left corner in front of a slide displaying ChatGPT and WolframAlpha logos. On screen: OUTPUT ERRORS ChatGPT WolframAlpha

27:26–28:03 The presenter stands in front of a white background with comparison bar charts showing metrics for Google, ChatGPT, and Stable Diffusion. On screen: HIGH COST $0.02-0.08 $0.02 10x $0.002 Google ChatGPT Stable Diffusion $0.02/1000 tokens $0.02/image SLOW INFERENCE TIME 5-20sec 10-20x 0.5 sec

28:03–28:25 A man speaks on the lower-left portion of the screen in front of a graphic comparing Stable Diffusion inference speeds. On screen: SLOW INFERENCE TIME Stable Diffusion Default 6.4sec Optimized 3x 2.1sec Q4 2022 Q1 2023

28:25–28:43 A man speaks to the camera in the bottom-left frame against a screenshot of the ChatGPT interface with explanatory text overlays. On screen: LACK OF LONG TERM MEMORY AI in Science and Medicine AI and the Future Google's transformer model Internet Platform Evolution ChatGPT's best uses First…

28:43–28:44 Speaker A talks to the camera on the left side of the screen beside news headlines regarding AI copyright and data access. On screen: DATA ACCESS & COPYRIGHT CEO Elon Musk pauses OpenAI chatbot's access to Twitter's database As the world experimented with OpenAI chatbot service, Twitter CEO…

28:45–28:53 Speaker A continues presenting while gesturing, framed beside news articles about AI copyright disputes and data access. On screen: DATA ACCESS & COPYRIGHT CEO Elon Musk pauses OpenAI chatbot's access to Twitter's database As the world experimented with OpenAI chatbot service, Twitter CEO…

28:53–29:19 A presenter in a patterned shirt speaks alongside a slide displaying news articles regarding AI data access and copyright lawsuits. On screen: DATA ACCESS & COPYRIGHT CEO Elon Musk pauses OpenAI chatbot's access to Twitter's database As the world experimented with OpenAI chatbot service, Twitter CEO…

29:19–29:51 A man speaks in the lower left against a graphic showing stages of data access and copyright with creative app logos. On screen: DATA ACCESS & COPYRIGHT RESEARCH CREATE DISTRIBUTE MONETIZE PREVIOUSLY dribbble Bēhance shutterstock

29:51–30:30 A man speaks facing the camera with a graphic on screen detailing creative workflow tools before and after AI. On screen: DATA ACCESS & COPYRIGHT RESEARCH CREATE DISTIRBUTE MONETIZE PREVIOUSLY dribbble Bēhance Canva shutterstock NOW ChatGPT Dall-E 2

30:30–30:37 A presenter in the lower left talks enthusiastically beside a slide comparing traditional creative workflow tools with generative AI tools. On screen: DATA ACCESS & COPYRIGHT RESEARCH CREATE DISTIRBUTE MONETIZE PREVIOUSLY dribbble Bēhance Canva shutterstock NOW ChatGPT Dall-E 2

30:41–30:49 A presenter speaks enthusiastically in front of an on-screen graphic comparing creative workflows before and after generative AI. On screen: DATA ACCESS & COPYRIGHT RESEARCH CREATE DISTIRBUTE MONETIZE PREVIOUSLY dribbble Bēhance Canva shutterstock NOW ChatGPT Dall-E 2

30:49–31:09 A man speaks on the left side of the frame while a graphic displaying a tweet from Elon Musk is shown on the right. On screen: CLOSED SYSTEMS & BIASES Elon Musk @elonmusk Replying to @GRDecter OpenAI was created as an open source (which is why I named it “Open” AI), non-profit company…

31:09–31:46 A man speaks on the left while historic illustrations of the Gutenberg press and early locomotive regulations appear on the right side of the screen. On screen: CLOSED SYSTEMS & BIASES LOCOMOTIVE ACT 1861 (SECTION 6.) NOTICE TO OWNERS AND DRIVERS OF LOCOMOTIVES THIS BRIDGE IS INSUFFICIENT TO CARRY WEIGHTS BEYOND THE…

31:46–32:00 A presenter speaks in the lower-left overlay against a background displaying the Stability AI homepage. On screen: stability.ai API News FAQ English CLOSED SYSTEMS & BIASES AI by the people, for the people Designing and implementing solutions using collective intelligence…

32:00–32:07 A presenter speaks in the foreground in front of a presentation slide showing a stage interview with Reid Hoffman. On screen: AI IS ABOUT TO RESET EVERYTHING

32:07–32:24 Sam Altman sits in an armchair against a blue circuit pattern backdrop while speaking. On screen: SAM SAM ALTMAN GRAYLOCK

32:24–32:29 Reid Hoffman and Sam Altman sit on stage discussing technology in front of a Greylock backdrop. On screen: greylock SAM ALTMAN GRAYLOCK

32:29–32:34 The host energetically explains economic reset by AI in front of a slide quoting Deloitte statistics. On screen: AI IS ABOUT TO RESET EVERYTHING For a typical Fortune 500 company, payroll is $1 to $2 billion per year, which averages between 50% to 60% of company spending…

32:34–32:47 The speaker addresses the camera while standing in front of a slide displaying Deloitte statistics regarding Fortune 500 payrolls alongside 3D rendered human avatars. On screen: AI IS ABOUT TO RESET EVERYTHING For a typical Fortune 500 company, payroll is $1 to $2 billion per year, which averages between 50% to 60% of company spending…

32:47–33:01 The presenter sits in the foreground talking to the camera beside a large graphic titled 'AI IS ABOUT TO RESET EVERYTHING' showing categories of future impact. On screen: AI IS ABOUT TO RESET EVERYTHING EDUCATION BUSINESS HEALTHCARE ART SOFTWARE GAMING

Transcript

0:00 Speaker A: We've definitely hit the Apollo landing moment where all eyes are on artificial intelligence and how it will reset everything. Speaker A: My AI buddy always nails the intros and I am incredibly excited to explore together with you how AI is going to change your and my life, why this is happening now and what are some of the coolest examples out there. Speaker A: If you asked somebody five years ago what would be the first skills to replace by artificial intelligence, they would have said the physical labor, warehouse workers, truck drivers, followed by the less demanding cognitive labor like accountants, then the high demand cognitive labor like computer development.

0:32 Speaker A: And lastly it's going to go for creativity because there's something so human about it that cannot be replaced by computers like creating art and films. Speaker A: Actually the exact opposite is happening at the moment where creativity is the first use case for artificial intelligence to go into masses, then followed by high demand in cognitive labor and others. Speaker A: And you can see it in the success of ChatGPT taking just less than two months to reach 100 million users compared to Facebook taking four years, or mobile's taking taking 12 years to do that.

1:00 Speaker A: It's one of the fastest platforms to grow, which is not entirely true as Pokemon go actually reached 100 million users in just less than a month. Speaker A: But that's a topic for another video. Speaker A: Artificial intelligence is top of mind for everybody around the world. Speaker A: That's seen on Google Trend. Speaker A: We're really entering in this new platform era. Speaker A: It started in the 1980s with PCs where we got browsers and search that we're still using. Speaker A: Building on top of that we got mobile and mobile really brought us those app stores where you can get any service as an app, as well as all the super apps and social media apps that we're using day to day.

1:33 Speaker A: That was building on top of the cloud platform which we've been using so far. Speaker A: That brought cross platform usage. Speaker A: It just didn't matter which device you're holding in your hand and collaboration in the cloud. Speaker A: Now we're entering this new phase in artificial intelligence which is bringing first intelligent agents and who knows what in the future that we're going to explore later. Speaker A: But why is this happening now? Speaker A: Let's hear from the co founder and CEO of OpenAI, Sam Altman.

1:58 Speaker B: We had the model for ChatGPT in the API for like, I don't know, 10 months or something before we made ChatGPT. Speaker B: And I sort of thought someone was going to just build it or whatever and that, you know, enough people had played around with it. Speaker B: Definitely if you make a really good user experience on top of Something like one thing that I very deeply believed was the way people wanted to interact with these models was via dialogue.

2:24 Speaker B: And we kept telling people this, we kept trying to convince people to build it and people wouldn't quite do it. Speaker B: So we finally said, all right, we're just going to do it. Speaker A: And they just did it. Speaker A: And really I think there's more than just the interface. Speaker A: But that's the key part because all the pieces were in place already 10 months before, as you said, but nobody was really using it. Speaker A: I believe the key drivers are the quality of the models themselves, interface, which we're going to talk first, and cost. Speaker A: So let's dive in on interface.

2:49 Speaker A: You can see on the left the ChatGPT, not ChatGPT or GPT3 playground where you could actually have a chat experience already one and a half years ago where all the models were pretty much in place with slight tweaks that were made to ChatGPT. Speaker A: So even the slightest changes in the interface and design and the model can make a huge difference. Speaker A: The interface for AI has really changed over time. Speaker A: We started with those base models that are still around like GPT3, but it's really hard to interact with them because you need to be really precise and a good prompt engineer to get any information out that you need.

3:21 Speaker A: Then we got those instruct models. Speaker A: That means the input output models. Speaker A: We had a specific goal in mind, for example to get weather information information or to get a quote from the book so the AI knew what to look for. Speaker A: And now more recently, just a couple of months ago, we've been seeing those exchange interfaces where you can have a two way conversation like ChatGPT. Speaker A: On top of that we now have this retrieval interface model. Speaker A: And that's something you see in the Bing Chat where you put that, take that chat GPT experience and put that together with the information, the live information that you can retrieve from the real life, so you can also quote the results.

3:58 Speaker A: And that's the stage we're at now. Speaker A: The ultimate version of the interface for AI that everybody's chasing is the action model. Speaker A: That's where everybody has your login details as well as your credit card information. Speaker A: So you can actually do it for you rather than you having to instruct it or have a conversation or try to retrieve information. Speaker A: Now you can just go straight to action and you see it being used everywhere. Speaker A: Like Replit in the code code generator that's built into the tool they're using.

4:23 Speaker A: Definitely like some of the models that we talked about with, or Edge that is using the retrieval model, putting together with a conversational model to optimize the results. Speaker A: So the interface is evolving. Speaker A: The second thing that we talked about, why it's happening in addition to interface is cost. Speaker A: We talked about in my previous video how the cost of making videos using AI is now almost nothing compared to just 10 years ago where you had to pay several hundred thousand for a camera and tools. Speaker A: If you want to see that video, by the way, click on the video up here and check it out.

4:51 Speaker A: How it's going to change Hollywood, where you can make videos on your laptop using, you know, GPT as a writer, using Descript or other voice platforms, as an audio, and using Synthesia for example, to create the video or any other video generation platform. Speaker A: The same thing is about to happen for intelligence. Speaker A: What it used to cost you quarter of a million dollar to get a degree in an Ivy League college. Speaker A: You could get now an online course, an online certificate for just, you know, $10,000.

5:17 Speaker A: Or you can just use the ChatGPT interface for $20 a month and have access to all of the intelligence of the world. Speaker A: The books that you can imagine, the website you could never visit. Speaker A: This is happening because, number one, we have so much data out there. Speaker A: Everything that is done on computers is now tracked and part of those models. Speaker A: So the amount of data that we could use to model is insane. Speaker A: At the same time, the cost of compute and energy has been coming down rapidly.

5:43 Speaker A: So building those models and using those models is becoming cheaper and cheaper. Speaker A: And that's why the first to go really are the works that are done on computer. Speaker A: They're using bits that can actually be used to model out rather than atoms, the physical elements, the warehouses, which are much harder to track as all of it is not yet connected to the Internet. Speaker A: Last but not least is the quality of the models. Speaker A: Why is AI now in a space or machine learning in a space where we can find it more useful than ever?

6:11 Speaker A: It really started in the 90s when the algorithmic models were developed. Speaker A: Those were the simple ways of digitizing. Speaker A: When we were starting to digitize our assets, like, you know, accounting, going to a computer so we can do simple calculations, search and filtering, input output methods. Speaker A: It then evolved into data science that was driven by big data, where everything started to be connected to the Internet and computers and we started to do more and more of our applications online and on computers.

6:36 Speaker A: Hence we started building those larger, more complex algorithms and using data science to resolve them. Speaker A: And then we evolved into machine learning, which is really a subset of artificial intelligence in the middle of 2010s where we had humans as well as then computers developing the algorithms. Speaker A: And that's where deep learning models were developed and where computers really became better in many things. Speaker A: For example, in understanding image classification, where they surpassed human accuracy in 2015, or even for classifying audio or replicating audio that was surpassed in the mid 2010s to now that has evolved into large generative AI models which have less human in the loop, which are using completely new methods of artificial intelligence and the transformers methods as well as the new attention method.

7:23 Speaker A: So it's not somebody guiding the AI to go from one step to another, but we really give free hands to the AIs to figure out themselves where the attention should be in the model and build multiple hundreds of versions of that. Speaker A: And that's why the quality has got so good. Speaker A: And these three are the key drivers for making AI paradigm shift. Speaker A: You know, the quality of the models, achieving human level intelligence, their interface and user experience turning into chat, human like conversations as well as the tools trending rapidly towards zero.

7:50 Speaker A: And that's what made ChatGPT success. Speaker A: Obviously we as humans have to evolve with it. Speaker A: We started on being, you know, those passive human beings who were foraging around gathering their own goods to laborers who started specializing in different skill set. Speaker A: And that gave a start to communities, cities and agriculture. Speaker A: To now the stage where we are at, where we started building machines to do the work for us. Speaker A: We started building computers and manufacturing machines. Speaker A: So we became the creators and operating those machines.

8:17 Speaker A: And the stage we're entering now is becoming narrators so that we don't build the machines, but we just describe what needs to be built and the machines will build themselves. Speaker A: And this does not just apply to language, which we talked about a lot. Speaker A: This also goes to image modeling, audio modeling, video modeling, code modeling and science modeling. Speaker A: We're going to dive into each one of them, how they're developed, and some of the coolest examples that you've probably not seen yet. Speaker A: So first of all, what are some of the largest language models out there?

8:42 Speaker A: Obviously ChatGPT we've talked about a lot of. Speaker A: In addition to that, we have Bing Chat that has evolved on top of that, to get the retrieval model out there. Speaker A: Google has announced part of their own chat version of Search as well as Anthropic which has developed a model called Claude. Speaker A: It's interesting that the talent has really been moving from Google to OpenAI to develop those models. Speaker A: And OpenAI people have now moved to Anthropic to develop new language models. Speaker A: They have Hundreds of billions of parameters.

9:08 Speaker A: Baidu has one of the largest models out there with 260 billion parameters. Speaker A: And many of the other companies that are not on this slide are also developing their own large language models. Speaker A: But how do you actually develop one? Speaker A: Let's use OpenAI as an example. Speaker A: First of all, you need a ton of data that includes in their case, all of the books that you can imagine. Speaker A: Substack, Reddit, Twitter, Wikipedia and Common Crawl, which really takes almost all of the website data and puts it in a structured form for anybody to access.

9:35 Speaker A: That's about 45 terabytes of data. Speaker A: And then all of that data is sent for training using the transformer infrastructure to the Nvidia's most powerful A100 chips. Speaker A: And all of that compute power will cost about 5 million to generate. Speaker A: And then from that, from that 45 terabytes of data and training, you get a GPT3 model which is about 800 gigabytes. Speaker A: So you could probably fit it on your computer, but to run it, it still cost quite a bit of money.

10:00 Speaker A: That's the inference part where you actually get the results. Speaker A: And it costs about anywhere between 2 to 8 cents depending on the length of the response that the OpenAI ChatGPT is providing for the server cost. Speaker A: Of course, there's many more models out there. Speaker A: If you want to give it, check them out. Speaker A: Go to LifeArchitect AI and he's also got a great YouTube channel that keeps tracks of all of the new models coming out and you can rank them based on different size. Speaker A: You will find a link for that down below in the description.

10:26 Speaker A: The uses for language models actually go beyond what people imagine today. Speaker A: And what is on this slide number one is of course text generation. Speaker A: We talked about that a lot where you give a prompt and you expect to get a response back. Speaker A: You could also use it for summarization. Speaker A: If you have a large PDF in front of you and you want to just answer your question and find that nugget that will be useful for you. Speaker A: You can use it for that, you can use it for translation. Speaker A: The language models are so much better than Google Translate ever really was and makes the knowledge accessible that is currently only locked to a couple of languages, to everybody in any language.

10:57 Speaker A: Of course you can use it for search and that's what Bing reimagines their own experience around. Speaker A: Getting more precise and more conversational answers straight in the search rather than having to search for them on the websites. Speaker A: And it becomes an intelligence assistant, helping you use applications that maybe you don't even know how to use the user interface, but you can use natural language to give the prompts and guidance. Speaker A: And last but not least, you can use it for analysis. Speaker A: So help to derive data from different websites or from the whole language model to help you make an easier purchase decision.

11:27 Speaker A: For example, an example of one of them is Jasper, which has built a platform that helps you to do any kind of writing that you can imagine straight on the platform. Speaker A: So let's see how it actually works in practice. Speaker C: Jasper is here to help. Speaker C: It's an app that uses AI to help you create any kind of content you need. Speaker C: It's like having your very own robot sidekick with an infinite supply of ideas. Speaker C: And that's why the perfect pair of socks is like a hug for your feet.

11:53 Speaker A: Damn, that's good. Speaker A: And you can use it really for everything from generating ideas to generating full scripts. Speaker A: And here's an example how Mint Mobile was using it right really in practice. Speaker A: So they gave prompts to OpenAI to use a joke curse word and big wireless and still going titles in the video. Speaker A: And then Ryan Reynolds was seeing if ChatGPT was doing its job and then reading it back and this becoming a real ad that they were running earlier this year.

12:18 Speaker D: First of all, let me just say Mint Mobile is the shit. Speaker D: But here's the thing. Speaker D: All the big wireless companies out there are ending their holiday promos, but not Mint Mobile. Speaker D: We're keeping the party going because we're just that damn good. Speaker D: Give Mint Mobile a try. Speaker D: And hey, as an added bonus, if you sign up now, you'll get to hear my voice every time you call customer service. Speaker D: Just kidding. Speaker D: That's. Speaker D: That's not really a thing. Speaker A: And that's something where it becomes really practical for them to yield content, create awareness by using somebody as a writer that cost them $20 a month.

12:51 Speaker A: Bing recently launched Bing Chat. Speaker A: As I said, this redesigns and reimagines search. Speaker A: So rather than just giving a three word keyword, which most of the people do and half of the people go straight back to search, you can actually give guidance and then go back and forth to find exactly what you need. Speaker A: This is to get, for example, craft ideas. Speaker A: The other example here is to build a fitness workout program. Speaker A: Again like for your needs exactly of what you're looking for and it's going to try and ask questions in return and then build the perfect program just for you and also referencing the websites where it's getting the information from.

13:22 Speaker A: So getting putting the large language model together with a retrieval model of what you can get from the websites and gives you the perfect answer. Speaker A: The second category is image modeling. Speaker A: Some of the largest ones out there that you might have heard of is obviously Dall E2, that you can go out there and use stable diffusion, which is open source. Speaker A: You can run it on your own computer and midjourney, which you can use in Discord in a chat environment. Speaker A: Google has built some of the strongest models out there as well, but they haven't really made them public yet.

13:48 Speaker A: But they are one of the largest in terms of parameters. Speaker A: How do you build an image model? Speaker A: Stable diffusion is a good example. Speaker A: Here you use the data. Speaker A: They have large open data like Lay on, which has got 5 billion text and image pairs. Speaker A: So you to put together the description of the image together with what the image looks like. Speaker A: They also had a deal with Getty Images which we're going to talk about later, A controversial deal. Speaker A: You put that 120 terabytes of data together and then based on the CEO of Stability AI, it took about 150,000 hours of compute for the most powerful Intel, a 100 Nvidia NY100 chip, and cost them about $600,000 to build the model.

14:27 Speaker A: And then the model itself is just a 2 gigabyte file. Speaker A: So getting from 120 terabytes just down to 2 gigabytes and you could run it right on your iPhone or right on your Mac or any computer that you imagine for free, using your own graphics card and your own compute. Speaker A: Or you could go out and obviously use any other service that you pay a few dollars for. Speaker A: What can it be used for? Speaker A: Text to image. Speaker A: Of course, we talked about that. Speaker A: Image to image. Speaker A: I'm sure you've seen Lens AI where everybody's sharing their avatars with those predefined prompts to go from image to image.

14:57 Speaker A: You can use it for outpainting. Speaker A: So you might have a small image, but you want to expand on that, fill in the space that's been done for many famous paintings out there. Speaker A: You can use it for inpainting, you know, to replace even somebody like Mona Lisa's hair with Mohawk. Speaker A: You can use it for enhancements and editing, for color correcting. Speaker A: Or you could also use it for classification and detection of different objects inside the images. Speaker A: So let's look at some of the examples out here. Speaker A: Now, first of all, Dall?

15:22 Speaker E: E that can take simple text descriptions, like a koala dunking a basketball, and turn them into photorealistic images that have never existed before. Speaker E: Dall? Speaker E: E 2 can also realistically edit and Retouch photos based on a simple natural language description. Speaker E: It can fill in or replace part of an image with AI generated imagery that blends seamlessly with the original. Speaker A: It's called inpainting and moreover, you could also use it not just for generating images, but from something practical point of view.

15:50 Speaker A: This is an example for designing an interface using using a tool called Galileo where you can say, hey, create me an onboarding screen for this dog walking app and then also create the settings page with username and photo and instantly you will have that generated by AI again using the database that they have. Speaker A: Or you then want to create the profile page for completely something else, another another app. Speaker A: And by the way, while you generate all of that, this becomes transferable that you can take it out for example to Figma and then from Figma launch it as an actual app.

16:22 Speaker A: So imagine how far you can take it. Speaker A: Audio is another big area of modeling, specifically in speech modeling and music, where Google has been building large language models like Audiolm which is building on WaveNet before Amazon has built Polly, which is the largest, one of the largest voice libraries. Speaker A: OpenAI has Whisper, the Text to Speech Library, Speechify and many others. Speaker A: And in music space we're seeing more recently developments like MusicLM to generate music from text or sing a song singsong which you can use audio to audio as well as one of the largest open source movement, Uber Duck.

16:55 Speaker A: What are some of the use cases? Speaker A: Of course number one is text to speech. Speaker A: You can just type something up and it be spoken back to you. Speaker A: The second thing is taking a library of audio files and turning them into text and maybe making them searchable. Speaker A: Using AI you can use speech to speech in different languages or even in the same language. Speaker A: You can just change your voice using something like re speech or you can do text to music. Speaker A: You can just keep prompts like this is the type of vibe I want, this is the type of style I want and it will generate music for you.

17:22 Speaker A: Or you can hum some voice and it will fill in the background like the voice to music model. Speaker A: Or you can just do enhancements for the audio, you know, make the audio sound better or exactly what you want it to be described using a text model. Speaker A: Here's an example how a company called Podcast AI brought back models from Joe Rogan Voice as well as Steve Jobs voice and put them in a podcast format to see what they would say if Steve was still alive.

17:48 Speaker F: It's always good to see you buddy. Speaker F: I'm so happy you came on man. Speaker A: Yeah, it's great to be on the show. Speaker A: Your audience is just so different from your normal Apple users. Speaker A: And that's a good thing. Speaker G: That's cool. Speaker F: Well, you know, I was an Apple user way before I did this show. Speaker F: I've been a fan of yours, macintosh, since the 1980s. Speaker A: I think it sounds great, not just from content point of view, but also from the sound of the voice. Speaker A: Of course it's going to get better and better over time.

18:14 Speaker A: And this is an example of somebody who just used ChatGPT to come up with rap rhymes for Eminem and then used Uber Duck to integrate make them into sounds and then integrate into them his song. Speaker A: So this is what you're gonna hear now. Speaker A: This is the future rave sound. Speaker A: I'm getting Austin and I'm in Brown. Speaker A: This is the future rave sound. Speaker A: I'm getting Austin and I'm in Brown.

18:43 Speaker A: Eminem, bro, there's something that I made as a joke and it works so good I could not believe it. Speaker B: I discovered those websites that are about. Speaker A: AI and this is how much you can do with just little experiments that we're starting to see happen around us. Speaker A: Moving on to video, and this has really two major categories. Speaker A: One is 3D category, which is generating completely new objects that can be editable in any format.

19:08 Speaker A: And the other one is avatars, which is somebody generating a lookalike like we did in the last video that you saw in 3D. Speaker A: Some of the largest models are Google's image and video. Speaker A: Of course, Meta showed their own make a video platform as well as Kirky AI, which we're going to see in example later in the avatar space. Speaker A: There's Synthesia, there's DID that can take photos to turn them into images and rephrase and many others that don't fit on this slide. Speaker A: What are some of the example use cases for video models?

19:35 Speaker A: Of course, you can use text to video. Speaker A: Just give it a prompt and it's going to generate video out of it. Speaker A: You can do image to video, just give it a photo and then turn that into a moving image. Speaker A: You can personalize video edit just, you know, some templated objects inside the video to make it very unique. Speaker A: You can do text to 3D. Speaker A: You can describe something and then turn that into a 3D object or a moving object. Speaker A: You can use it for classification, for example, for fitness apps to see how well you're tracking against the moves.

20:02 Speaker A: And you can use it for enhancements and editing up resolutions, down resolutions, clarifying, etc. Speaker A: So let's see some of the examples and the first one here is the Runway ML and their new model to stylize any content, generate video in any. Speaker F: Style, all while retaining quality and flexibility. Speaker F: Gen1 is able to realistically and consistently apply the composition and style of an image or text prompt to the target video, allowing you to generate new video content using an existing video.

20:32 Speaker F: We call this approach Video to video. Speaker A: In addition to video video, we also talked about photo to video. Speaker A: And a good example Here is the DID platform. Speaker H: DID's Creative Reality Studio just got even better. Speaker H: Now as well as animating images from text or audio, you can create avatars and generate scripts using stable diffusion and GPT3. Speaker H: Just describe what you want your avatar to look like and press Create. Speaker H: You can make any avatar you want, say picture of a half alien goddess or a 19th century portrait of a happy Victorian gentleman.

21:03 Speaker A: And then the other example here, going beyond just avatars and changing the style of the movies, is also to generate full 3D objects. Speaker A: An example here is the Kirky AI where you can use text to 3D prop.

21:31 Speaker A: And somebody's using it in real life. Speaker A: Already is Roblox where they're using generative AI and prompt guidance to design a game. Speaker A: Of course they've got a whole host of tools out there that you can use today, but they're still pretty complex and only fraction of the users use them. Speaker A: So here you can just do text guidance of what do you want the car look like, what do you want the paint look like, what do you want the lights to do? Speaker A: And that automatically generates the code for you within the game that you can then launch on their platform, which makes it so easy and so much more accessible for people using natural language also in game design or potentially even making movies inside games.

22:06 Speaker A: Next one is code and as you saw a little example in Roblox, it takes way beyond that. Speaker A: Some of the largest models out there are Codex by OpenAI as well as Copilot by GitHub which is built on top of Codex. Speaker A: Really? Speaker A: And Replit was built Ghostwriter and Google, which in secret is about to announce their own coding tool which they're developing internally, but there's no outside information available yet. Speaker A: What do you use the code models for?

22:31 Speaker A: Of course, number one is auto filling, helping you get your more done faster. Speaker A: The other one is text to code, so you can describe something and it will turn into code like we saw with Roblox. Speaker A: Or you can do Q and A and reviews, finding the errors in the code. Speaker A: Or you can do code to text. Speaker A: Somebody sends you a code but you don't Understand what's going on, you can get it translated into human language. Speaker A: Here's an example from Replit. Speaker A: In the first one you can see, hey, what does this index file do? Speaker A: And it tells you back in natural language of what the file is about to do.

22:59 Speaker A: Then it defines an error. Speaker A: Hey, the error is on line six and here's how to fix it. Speaker A: So you can copy the code, put it back to your editor and it fixes the result for you. Speaker A: And then you could use the chat interface to even develop further a new code for you, or even run QA on your existing code. Speaker A: This is how powerful it is. Speaker A: And Copilot has shared some results so that it makes developers 55% faster and almost half of the code that they do is now written by artificial intelligence.

23:27 Speaker A: Last but not least, let's talk about science, because that's where some of the biggest breakthroughs might come through. Speaker A: DeepMind and Google have built AlphaFold 2, which is one of the most critical protein folding prediction models out there. Speaker A: That is also something that Meta is working on with ESM Fold as well as many of the other research bodies, including some of the open source ones like the OpenBioML, who is building together with other partners Librefold Project. Speaker A: What are some of the uses for science models? Speaker A: Number one is review and analysis.

23:53 Speaker A: You can take in all of the science papers and then understand what are the correlations, what others have been working on that humans can't possibly work through. Speaker A: Hundreds of thousands of papers that are coming out as they're exponentially growing and there's more and more every day. Speaker A: You can use it for predictive modeling, all of the phases, for example of the development of new pharmaceuticals. Speaker A: You can use it for precision medicine, matching the specific individual and their biology to the specific pharmaceutical.

24:19 Speaker A: You can use it for genomics, in really predicting how different proteins react. Speaker A: You can use it for diagnosis. Speaker A: So for example, taking a image of your retinal scanned and then matching it to a diabetics database to be much more accessible and more predictive in your diseases, as well as for monitoring, including in hospital and at home care, where you can use connected devices and the behavior models that you can build on top of that to understand if the patients are progressing as well as they could.

24:46 Speaker A: And here's a quick video to explain how useful and valuable that is from Alpha. Speaker G: So we created an AI system to solve this problem called AlphaFold. Speaker G: It's trained on the sequences and structures of about 100,000 proteins painstakingly mapped out by scientists around the world today. Speaker G: It can accurately Predict a protein shape just from its sequence of amino acids. Speaker G: Alphafold's predictions could enable progress in all sorts of areas.

25:13 Speaker G: Imagine a future where we can understand diseases more quickly and develop drugs to fight them, or one where we could use enzymes to break down plastic waste or even to capture carbon from the atmosphere. Speaker A: But it's not all rosy. Speaker A: For AI and generative AI, there are several challenges, including the hallucinations and the output errors, the high cost of the models, the lack of long term memory, the slow inference time and data access, as well as the biases in the system.

25:38 Speaker A: And let's go through each one of them and see what are the potential solutions out there. Speaker A: You have lot of articles that show you how the AI models are lying or they're building a relationship, they're in love, they're sentient. Speaker A: But actually I don't believe that really to be the case. Speaker A: I believe that's very much the kind of reaction the monkey would have with a mirror, you know, trying to fight off somebody that they see, that they think is somebody else, but it's actually themselves. Speaker A: And I feel we're in a very similar spot with prompts at the moment.

26:05 Speaker A: We are prompting the models to say the weird things back to us, we're prompting the model to go into a mode where he would kind of act back to us and what we're expecting from it. Speaker A: And I think this is exactly what's happening today. Speaker A: It's more like a monkey in a mirror effect than really the AI models being sentient because they are still need to be prompted. Speaker A: They wouldn't do anything by themselves or on their own. Speaker A: And of course there are already now more guardrails in place from for example, Bing, who doesn't even want to talk about it.

26:31 Speaker A: When you ask if it's alive, if it's sentient, does he love you? Speaker A: All of those are now blocked. Speaker A: And it just says it's something that they prefer not to talk about. Speaker A: And the output errors, that's another part of hallucinations is we saw those errors in the Google demo when they announced Bard. Speaker A: But if you go to Reddit subreddits like ChatGPT or Bing, you'll see a whole lot more on how people are trying to trick the systems. Speaker A: And I think that's very normal to push the boundaries and make them better.

26:58 Speaker A: And today the generated models are not great being factually correct. Speaker A: But I think in the phase of where Bing is at, where they're putting the large language model together with the exchange model, together with the retrieval model, that's where you can really get good facts. Speaker A: And that's an opportunity for many other companies like for example wolframalpha to put the mathematics models and calculations together with the chat models of ChatGPT and Make Something that is factually correct but also is something that is more useful and easier for consumers to use in a natural language way.

27:28 Speaker A: The second thing is high cost. Speaker A: It costs almost 10 times more from just a fraction of a cent on Google search to about 2 to 8 cents to run a query on ChatGPT. Speaker A: Of course, doing the same with images today in stable diffusion also cost you about 2 cents if you're buying it on their own servers, running it rather than running it on your own. Speaker A: But I'm sure that's something that's going to change in the future as the compute gets again cheaper and they get more powerful.

27:54 Speaker A: The slow inference time is directly connected to the cost. Speaker A: While it takes half a second to search something on Google, it takes about 5 to 20 seconds to get a prompt back in ChatGPT. Speaker A: Of course I believe that that's also going to change. Speaker A: Just over the last six months, stable diffusion has managed to decrease and increase the efficiency of the model by about three times by not really getting any reductions in the image quality. Speaker A: So inst of taking 6 seconds, it now takes 2 seconds to generate an image.

28:22 Speaker A: And I think the same thing is going to apply to cost and speed in the future. Speaker A: The fourth area is the lack of long term memory today. Speaker A: Similarly, the ChatGPT model can only memorize 4,000 tokens which is about 3,000 words in human language. Speaker A: And that's obviously limiting factor if we want those assistants to be personal for everybody. Speaker A: But I'm sure again that's going to change as the compute and the models get more efficient data access and copyright.

28:47 Speaker A: That's a key area. Speaker A: We've seen so many copyright issues now with, you know, stable diffusion being sued by Get Images as they integrate those watermarks and they clearly see that their database has been used to generate the large image models. Speaker A: You can see Elon Musk finding out that OpenAI has turned from a nonprofit to a for profit organization and they've closed the Twitter access, Twitter data access for OpenAI as well as many other artist groups and developer groups suing the large language models developers for the access of their data.

29:19 Speaker A: But I think it's not really that dissimilar of what we used to do before. Speaker A: You know, if imagine if I'm a designer or photographer, I might go to Dribbble and Instagram and Pinterest and Behance to get all the inspiration and ideas in place. Speaker A: I will then use, you know, Canva and Figma as a tool to create it. Speaker A: I will then put it on Behance, for example, as part of my portfolio and then I'm going to go and sell it on Shutterstock. Speaker A: I believe that if I then use, you know, copyrighted material, of course I'm going to be penalized and I shouldn't be able to make money with other people's content.

29:47 Speaker A: That's the same in this flow. Speaker A: Even though I took inspiration from them in my research phase. Speaker A: This is now just being done on superpowers, you know, rather than me being able to go maybe about 50 to 100 pages for inspiration. Speaker A: ChatGPT has done billions of pages for my inspiration. Speaker A: So I can use it as an input to get new ideas and develop new concepts. Speaker A: I can then use a tool like Dall Etude or Stable Diffusion or anything else to create that content.

30:13 Speaker A: I will then put it in my portfolio and then we'll monetize it in Shutterstock. Speaker A: By the way, Shutterstock has got a partnership with OpenAI already to start monetizing AI generated images. Speaker A: Because I believe that this is just a super powerful research tool bringing you a lot more in a more simpler way. Speaker A: And it's a super powerful creation tool, accelerating the process of creating and making it a lot more accessible for anybody. Speaker A: But the copyright really comes in when you start monetizing it.

30:39 Speaker A: So I wouldn't really see a version of where you can penalize those models as they're just there for inspiration. Speaker A: That's what we do every day. Speaker A: Artists and musicians take inspiration from others all the time. Speaker A: Now we can just do it on steroids. Speaker A: Last but not least, these are, you know, getting to be closed systems. Speaker A: You can see how OpenAI turned from open source and that's something that Elon Musk is really concerned about as well to something that is now just serving Microsoft. Speaker A: And it was really created as an anti Google move to be something that is open as it will be dangerous in his mind if it gets locked to just one organization.

31:10 Speaker A: But it's never been really different. Speaker A: Take Gutenberg Press. Speaker A: When that was announced, churches went in and they actually created licensing for every Gutenberg Press to be out there. Speaker A: What can you print? Speaker A: How many can you print? Speaker A: Where can you print? Speaker A: They were all regulated by church. Speaker A: Or take even the Locomotive act in the middle of 19th century in the UK just to protect the horses and the industry that was built around it. Speaker A: There was a law where it had to have somebody walking in front of the car going only four miles an hour maximum and waving a red flag as they were doing it.

31:42 Speaker A: Those were rules to protect and build a closed system. Speaker A: But the market has figured it out. Speaker A: In the AI space, there's a company called Stability AI who's behind Stable Diffusion, who is building AI by the people and for the people, aiming to make it open source and accessible for everybody, for any country and any individual to build their own models. Speaker A: Overall, AI will and is about to reset absolutely everything for the two reasons.

32:08 Speaker B: My basic model of the next decade is that the cost of intelligence, the marginal cost of intelligence and the marginal cost of energy are going to trend rapidly towards zero, like surprisingly far. Speaker B: And those I think are two of the major inputs into the cost of everything else except the cost of things. Speaker B: We want to be expensive, the status, goods, whatever. Speaker A: And that's why everything will about to be reset by AI, because intelligence and energy will be marginal.

32:35 Speaker A: That's at the moment about 50 to 60% of the cost of the Fortune 500 companies is the payroll paying for the labor. Speaker A: And if the labor and the intelligence will become essentially free for everybody, then that's going to rewrite how we do, how we get education, how it affects businesses, healthcare, art, software and gaming, which we're going to talk about in the next video. Speaker A: So thank you very much for watching. Speaker A: I hope you enjoyed it. Speaker A: Let me know what you want to hear about first and see you next time.