Tag Archives: llms

Yann LeCun’s $1B Bet Against LLMs

Video

Video Description

Apply to join Hudson River Trading: https://www.hudsonrivertrading.com/welchlabs
Welch Labs Book: https://www.welchlabs.com/resources/ai-book-ezrzm-msrmc
Patreon: https://www.patreon.com/c/welchlabs

Sections
0:00 – Intro
2:28 – The Problem with Deep Learning
4:17 – Intelligence is a Cake
5:15 – The Rise of Generative AI
8:00 – Blurry Images
8:54 – HRT is an awesome place to work
11:16 – But why so Blurry?
13:30 – Do our models need to be generative?
15:16 – Siamese Networks
17:53 – Representation Collapse
19:54 – Yann’s Epiphany u0026 Barlow Twins
27:22 – DINO
28:58 – JEPA u0026 World Models
34:09 – But is JEPA good?
36:19 – Welch Labs Book

Special thanks to: Yann LeCun, Stephane Deny, David Fan, Nicolas Ballas

Clip of Yann from 1989: https://www.youtube.com/watch?v=FwFduRA_L6Q

CNN Paper: http://yann.lecun.com/exdb/publis/pdf/lecun-89e.pdf
LeNet-5 paper: http://vision.stanford.edu/cs598_spring07/papers/Lecun98.pdf

Dashcam video
https://commons.wikimedia.org/wiki/File:Car_Driving_Faadou_4K_HDR-_Rural_road_-_Canton_-_327.webm

Image Credits
https://en.wikipedia.org/wiki/File:Dota_2_Gameplay_Aug_2017.jpg
https://commons.wikimedia.org/wiki/File:Felis_catus-cat_on_snow.jpg
https://commons.wikimedia.org/wiki/File:Magnificent_CME_Erupts_on_the_Sun_-_August_31.jpg
https://commons.wikimedia.org/wiki/File:Alcedo_atthis_-_Riserve_naturali_e_aree_contigue_della_fascia_fluviale_del_Po.jpg
https://commons.wikimedia.org/wiki/File:Biandintz_eta_zaldiak_-_modified2.jpg

V-JEPA2 Robot Arm Videos
https://ai.meta.com/research/vjepa/

PATRONS
Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely, Jeff Eastman, Rodolfo Ibarra, Clark Barrus, Rob Napier, Andrew White, Richard B Johnston, abhiteja mandava, Burt Humburg, Kevin Mitchell, Daniel Sanchez, Ferdie Wang, Tripp Hill, Richard Harbaugh Jr, Prasad Raje, Kalle Aaltonen, Midori Switch Hound, Zach Wilson, Chris Seltzer, Ven Popov, Hunter Nelson, Amit Bueno, Scott Olsen, Johan Rimez, Shehryar Saroya, Tyler Christensen, Beckett Madden-Woods, Darrell Thomas, Javier Soto, U007D, Caleb Begly, Rick Rubenstein, Brent Hunsaker, Dan Patterson, Tchsurvives, Alex Adai, Walter Reade, Zyansheep, Walter Reade, Duncan Stannett, Reginald Carey, Jean-Manuel Izaret, dh71633, Adrian Rodriguez, Dimitar Stojanovski, Michael Harder, Peter Maldonado, Emily Pesce, David Johnston, Insang Song, FaeTheWolf, Stephen Taylor, KittenKaboodle, EMatter, PATRICKMCCORMACK, John Beahan, Cameron, Cole Jones, Garrett Thornburg, Jeroen W, Rohit Sharma, GlennB, Emmanuel Cortes, Katie Quinn, Karina C, Cakra WW, Mike Ton, Eric Gometz, MacCallister Higgins, Niko Drossos, David Eraso, Tom Zehle, Steve, Brian Lineburg, rjbl, Michael Loh, Perry Vais, Bengal0, Farhad Manjoo, Sara Chipps, Ellis Driscoll, William Taysom, Will Harmon, CK, Abdullah, Peter Cho, Leo Nikora, Griffin Smith, Ash Katnoria, Alex, Markus Hays Nielsen, Catherine H., Vi, David Dobáš, Peter Wang, Sina Sohangir, Danny Thomas, Julian Francis, Hans Adler, Jiayu Peng, Weston M, Youssouf da Silva, John Thomas, Samuel Costello, Sam Adams, Bryan Liles, Malaya Zemlya, Karl, Vahe Andonians, Mike Doughty, Larry Novelo, Jonas Acres, Ludicrum Rex, Robert Blumofe, Anthony Z, Alex Zhao, Dan Babitch, Nikko Patten

Supporting code: https://github.com/WelchLabs/videos

Created by: Sam Baskin, Pranav Gundu, and Stephen Welch
Content ID: CFAQJOTYQHT7JYIT

LLMs and AI Agents: Transforming Unstructured Data

Video

Video Description

Read more about Terzo here → https://ibm.biz/Bdnmpr

Learn more about Intelligent Data Extraction here → https://ibm.biz/BdnqSX

Are AI agents the new assembly line for data? 🤖 Join Eric Pritchett from Terzo as he explores how LLMs, GPT models, and AI agents turn unstructured data into actionable insights. Discover how OCR, NLP, and agentic workflows reshape document intelligence and solve real-world challenges! ✨

Discover more about Terzo → https://ibm.biz/Bdnmps

#ai #llm #unstructureddata #aiagents

Transformers (how LLMs work) explained visually | DL5

Video

Video Description

Breaking down how Large Language Models work
Instead of sponsored ad reads, these lessons are funded directly by viewers: https://3b1b.co/support

Here are a few other relevant resources

Build a GPT from scratch, by Andrej Karpathy
https://youtu.be/kCc8FmEb1nY

If you want a conceptual understanding of language models from the ground up, @vcubingx just started a short series of videos on the topic:
https://youtu.be/1il-s4mgNdI?si=XaVxj6bsdy3VkgEX

If you're interested in the herculean task of interpreting what these large networks might actually be doing, the Transformer Circuits posts by Anthropic are great. In particular, it was only after reading one of these that I started thinking of the combination of the value and output matrices as being a combined low-rank map from the embedding space to itself, which, at least in my mind, made things much clearer than other sources.
https://transformer-circuits.pub/2021/framework/index.html

History of language models by Brit Cruise, @ArtOfTheProblem
https://youtu.be/OFS90-FX6pg

An early paper on how directions in embedding spaces have meaning:
https://arxiv.org/pdf/1301.3781.pdf

Звуковая дорожка на русском языке: Влад Бурмистров.

Timestamps

0:00 – Predict, sample, repeat
3:03 – Inside a transformer
6:36 – Chapter layout
7:20 – The premise of Deep Learning
12:27 – Word embeddings
18:25 – Embeddings beyond words
20:22 – Unembedding
22:22 – Softmax with temperature
26:03 – Up next

Deep Dive into LLMs like ChatGPT

Video

Video Description

This is a general audience deep dive into the Large Language Model (LLM) AI technology that powers ChatGPT and related products. It is covers the full training stack of how the models are developed, along with mental models of how to think about their "psychology", and how to get the best use them in practical applications. I have one "Intro to LLMs" video already from ~year ago, but that is just a re-recording of a random talk, so I wanted to loop around and do a lot more comprehensive version.

Instructor
Andrej was a founding member at OpenAI (2015) and then Sr. Director of AI at Tesla (2017-2022), and is now a founder at Eureka Labs, which is building an AI-native school. His goal in this video is to raise knowledge and understanding of the state of the art in AI, and empower people to effectively use the latest and greatest in their work.
Find more at https://karpathy.ai/ and https://x.com/karpathy

Chapters
00:00:00 introduction
00:01:00 pretraining data (internet)
00:07:47 tokenization
00:14:27 neural network I/O
00:20:11 neural network internals
00:26:01 inference
00:31:09 GPT-2: training and inference
00:42:52 Llama 3.1 base model inference
00:59:23 pretraining to post-training
01:01:06 post-training data (conversations)
01:20:32 hallucinations, tool use, knowledge/working memory
01:41:46 knowledge of self
01:46:56 models need tokens to think
02:01:11 tokenization revisited: models struggle with spelling
02:04:53 jagged intelligence
02:07:28 supervised finetuning to reinforcement learning
02:14:42 reinforcement learning
02:27:47 DeepSeek-R1
02:42:07 AlphaGo
02:48:26 reinforcement learning from human feedback (RLHF)
03:09:39 preview of things to come
03:15:15 keeping track of LLMs
03:18:34 where to find LLMs
03:21:46 grand summary

Links
– ChatGPT https://chatgpt.com/
– FineWeb (pretraining dataset): https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1
– Tiktokenizer: https://tiktokenizer.vercel.app/
– Transformer Neural Net 3D visualizer: https://bbycroft.net/llm
– llm.c Let's Reproduce GPT-2 https://github.com/karpathy/llm.c/discussions/677
– Llama 3 paper from Meta: https://arxiv.org/abs/2407.21783
– Hyperbolic, for inference of base model: https://app.hyperbolic.xyz/
– InstructGPT paper on SFT: https://arxiv.org/abs/2203.02155
– HuggingFace inference playground: https://huggingface.co/spaces/huggingface/inference-playground
– DeepSeek-R1 paper: https://arxiv.org/abs/2501.12948
– TogetherAI Playground for open model inference: https://api.together.xyz/playground
– AlphaGo paper (PDF): https://discovery.ucl.ac.uk/id/eprint/10045895/1/agz_unformatted_nature.pdf
– AlphaGo Move 37 video: https://www.youtube.com/watch?v=HT-UZkiOLv8
– LM Arena for model rankings: https://lmarena.ai/
– AI News Newsletter: https://buttondown.com/ainews
– LMStudio for local inference https://lmstudio.ai/

– The visualization UI I was using in the video: https://excalidraw.com/
– The specific file of Excalidraw we built up: https://drive.google.com/file/d/1EZh5hNDzxMMy05uLhVryk061QYQGTxiN/view?usp=sharing
– Discord channel for Eureka Labs and this video: https://discord.gg/3zy8kqD9Cp