Machine Learning Street Talk (MLST)
Machine Learning Street Talk (MLST)

Nieuwste aflevering
271 afleveringen
- A lot of people in AI treat evolution as a dumb fallback, basically random search for when you can't take a gradient. Akarsh Kumar thinks that is wrong. Selection hangs on to partial solutions, so mutations only need to be useful about 1% of the time for the search to keep making progress.Akarsh is a PhD student at MIT working with Phillip Isola, works with Sakana AI, and is first author of the Fractured Entangled Representation paper with Kenneth Stanley, Jeff Clune and Joel Lehman. He tells Tim Scarfe why the path a learner takes may shape the structure of what it learns, and why that is a different way of looking at intelligence from the statistical one.Most of the conversation is about ASAL, the method he led for searching whole spaces of artificial worlds. Rather than predict what a rule will do, ASAL runs the simulation and asks a foundation model what happened. Mapped across all 262,144 Life-like rules, the most interesting worlds sit on one small island. Along the way: Game of Life, Lenia (with a clip from its creator, Bert Chan), neural cellular automata, Boids, Particle Life and the emergence of persistence.---TIMESTAMPS:00:00:00 Intro: artificial life, ASAL and Core War in four minutes00:04:19 Life as it could be, not just as it is00:05:46 Why the order you learn things in matters00:08:33 No shortcuts: Wolfram, novelty search and regularisation00:10:44 Kenneth Stanley: the path matters, not just the destination00:11:31 Statistical intelligence vs regularity-based intelligence00:13:49 Artificial chemistry and other possible universes00:15:59 Convergent patterns: shadows of the substrate?00:18:02 How Conway's Game of Life works00:20:49 Change one cell, change everything?00:22:31 Lenia (with Bert Chan) and neural cellular automata00:25:26 Boids, Particle Life and cell-like creatures00:28:47 Persistence, entropy and what life is00:30:15 ASAL: a foundation model as the critic00:34:04 Impostors inside the simulation00:36:50 From primordial soup to alien animals00:38:26 262,144 rules and the island of open-endedness00:40:17 From artificial life to AGI00:41:35 Core War: Game of Thrones, Turing edition00:43:22 LLMs as the mutation step: evolving warriors00:46:03 Evolution is anything but random---REFERENCES:paper:[00:04:19] ASAL (Kumar et al.)https://arxiv.org/abs/2412.17799[00:08:08] Assembly theory https://www.nature.com/articles/s41586-023-06600-9[00:09:54] FEP paperhttps://arxiv.org/abs/2505.11581[00:22:38] Lenia: Biology of Artificial Life (Bert Chan)https://arxiv.org/abs/1812.05433[00:23:20] Growing Neural Cellular Automata (Mordvintsev et al.)https://distill.pub/2020/growing-ca/[00:41:35] Digital Red Queen: Core War with LLMs (Kumar et al.)https://arxiv.org/abs/2601.03335[00:43:42] MAP-Elites https://arxiv.org/abs/1504.04909[00:44:55] AlphaEvolve https://arxiv.org/abs/2506.13131[00:47:03] AutoML-Zero (Real et al.)https://arxiv.org/abs/2003.03384book:[00:07:04] Why Greatness Cannot Be Planned (Stanley and Lehman)https://link.springer.com/book/10.1007/978-3-319-15524-1tool:[00:18:12] Conway's Game of Lifehttps://en.wikipedia.org/wiki/Conway%27s_Game_of_Life[00:25:27] Boids (Craig Reynolds)https://www.red3d.com/cwr/boids/[00:27:12] Particle Life (Tom Mohr)https://github.com/tom-mohr/particle-life[00:41:50] Core Warhttps://corewar.co.uk/other:[00:08:49] Computational irreducibility (Stephen Wolfram)https://www.wolframscience.com/nks/p737--computational-irreducibility/[00:10:46] MLST: Why Every AI Model Is an Impostor (Kenneth Stanley, FER documentary)https://www.youtube.com/watch?v=o1q6Hhz0MAg[00:28:30] MLST: Blaise Agüera y Arcas on life emerging from codehttps://www.youtube.com/watch?v=rMSEqJ_4EBk---LINKS:Akarsh Kumar: https://akarshkumar.com/ASAL project page and demos: https://pub.sakana.ai/asal/Digital Red Queen project page: https://sakana.ai/drq/RESCRIPT:https://app.rescript.info/public/share/Qfv3T0EVzqOXL_CYYeRr9Blv8HnNm4JsXBh7ByFjJbc
- Tsung-Hsien (Shawn) Wen, CTO of PolyAI, tells Tim Scarfe why voice agents are harder than text agents. Voice adds time, and a good conversation depends on adapting to the person on the line, not just on reasoning to the best answer. Shawn describes an audio-native model (Dialog-RSN-1) that first predicts a turn-taking signal, then replies in text with citations, and writes the transcript last so enterprises can audit it.Along the way: training on real, noisy calls with synthetic noise added, and why over-cleaned audio made the new model worse. Latency, and what a voice agent should do while it thinks. Why a voice with a hint of regional accent beats a generic one. Why public benchmarks fall short for voice, why enterprises want to own their agent harness, and whether behaviour belongs in the harness or in the weights.The last stretch is about working with agents: cognitive debt, the shift from producing content to checking it, Wispr Flow, building tools that agents can use, and whether slop is in the eye of the reader.This episode was produced in partnership with PolyAI.https://poly.ai CHAPTERS0:00 Why voice agents are harder than text4:27 What enterprises want, and why PolyAI built its own model9:04 How an audio-native voice model works15:21 Training data, spectrograms and synthetic noise20:36 The cocktail party problem and the future of turn-taking25:21 Latency, adaptive reasoning and keeping callers' trust31:33 Voices, personality and the uncanny valley36:32 How do you benchmark a voice agent?42:00 Harness engineering and owning the intelligence45:09 Well-specified problems and auditable agents49:46 Weight adaptation and cognitive debt56:31 Agents at work: Wispr Flow, voice and tool building1:02:50 The next decade of voice, and what counts as slopREFERENCESThe Bitter Lesson: http://www.incompleteideas.net/IncIdeas/BitterLesson.html [9:05]Retrieval-augmented generation: https://arxiv.org/abs/2005.11401 [13:23]Mel scale: https://en.wikipedia.org/wiki/Mel_scale [17:16]Victor Zue: https://en.wikipedia.org/wiki/Victor_Zue [17:45]Cocktail party effect: https://en.wikipedia.org/wiki/Cocktail_party_effect [20:39]Speaker diarisation: https://en.wikipedia.org/wiki/Speaker_diarisation [21:16]
- Leonardo de Moura created Lean and co-created Z3. ---This episode is sponsored by Parallel.Parallel, where agents find answers: web search, extraction and deep research APIs built for AI agents.Start free with the Parallel MCP server and $5 of credits every month: https://parallel.ai/mlst?utm_source=creator&utm_medium=podcast&utm_content=MLST---Tim Scarfe talks with Leo about how Lean escaped its original audience, why dependent types and Mathlib made it useful to working mathematicians, and what happens when formal verification leaves the lab. De Moura explains the small trusted kernel and independent checkers, and gives his account of the recent Collatz incident, in which a purported proof was accepted by both Lean's official kernel and Nanoda, apparently by exploiting a different bug in each.---TIMESTAMPS:00:00:00 Cold open: the green checkmark can lie00:00:59 Cathedral or bazaar: who controls Lean's core?00:04:55 Why Lean's core stays small and protected00:08:07 The Slack purge, Brandolini's law and the Lean FRO00:11:12 The Collatz exploit: two kernels, two bugs00:16:44 More kernels, reward hacking and safety by transparency00:21:04 Sponsor: Parallel00:21:59 Kim Morrison, Claude and the zlib proof00:25:25 Can we specify complex systems?00:28:04 Specs change: proofs are cheaper to redo with AI00:31:27 From Lean 1 to Lean 400:34:57 Dependent types in plain terms00:37:25 Lean 4's extensibility and Mathlib's growth00:41:44 Mathlib as infrastructure: Formal Frontiers00:44:15 Creativity, abstraction and nut-sniping00:48:46 Breadcrumbs, not learning: what AI agents lack00:53:03 Competence without comprehension, and verified guardrails00:56:21 Is the human still the author?01:01:44 AlphaProof, LLMs and why certificates still matter01:06:38 What's next for Lean, and its legacy01:11:42 How to start learning Lean---REFERENCES:tool:[00:00:48] Leanhttps://lean-lang.org/[00:01:24] Mathlibhttps://github.com/leanprover-community/mathlib4[00:12:09] nanoda_libhttps://github.com/ammkrn/nanoda_lib[00:12:19] CollatzLeanhttps://github.com/xrchz/CollatzLean/blob/a79357462a33d2a6babd4cf6c8d8bcd25425d653/README.md[00:13:06] Lean issue 14576https://github.com/leanprover/lean4/issues/14576[00:13:21] Lean pull request 14577https://github.com/leanprover/lean4/pull/14577[00:13:45] nanoda_lib pull request 22https://github.com/ammkrn/nanoda_lib/pull/22[00:15:33] Lean comparatorhttps://github.com/leanprover/comparator[00:18:02] Lean4Leanhttps://github.com/digama0/lean4lean[00:19:50] ARC-AGI-3https://arcprize.org/arc-agi/3[00:21:59] lean-ziphttps://github.com/kim-em/lean-zip[00:29:45] CompCerthttps://compcert.org/[00:29:45] seL4https://www.sel4.org/[00:30:23] Z3https://github.com/Z3Prover/z3[00:38:10] Veilhttps://github.com/verse-lab/veil[00:38:10] Velvethttps://github.com/verse-lab/velvetorganization:[00:00:52] Lean FROhttps://lean-lang.org/fro/[00:42:39] Mathlib Initiativehttps://mathlib-initiative.org/about/person:[00:05:11] Ilya Sergeyhttps://ilyasergey.net/[00:10:31] Joachim Breitnerhttps://www.joachim-breitner.de/[00:32:58] Adam Chlipalahttps://adam.chlipala.net/[00:38:42] Kevin Buzzardhttps://www.ma.imperial.ac.uk/~buzzard/[01:10:16] Terence Taohttps://terrytao.wordpress.com/book:[00:08:19] The Proof in the Codehttps://us.macmillan.com/books/9780374620059/theproofinthecode/other:[00:10:05] Brandolini's lawhttps://en.wikipedia.org/wiki/Brandolini%27s_law[00:45:33] A new result on unit distanceshttps://openai.com/index/model-disproves-discrete-geometry-conjecture/[01:01:08] Fermat's Last Theorem formalisationhttps://imperialcollegelondon.github.io/FLT/paper:[00:34:36] The Lean 4 theorem prover and programming languagehttps://doi.org/10.1007/978-3-030-79876-5_37---RESCRIPT:https://app.rescript.info/share/7d3d4a0059443236a01f6c9acbf4db58https://app.rescript.info/api/public/sessions/b007264c0ce89047/pdf
- Weco let an AI coding agent rewrite the harness around another agent for eight days: its code, prompts and tools, while the underlying language model stayed fixed. Tim Scarfe asks Weco co-founder Zhengyao Jiang what the reported gains over two years of human engineering actually demonstrate.The discussion examines AIDE 85's generated code, held-out evaluation and the difficulty of separating useful discoveries from reward hacking. Jiang explains Weco's four levels of recursive self-improvement and compares the experiment with AlphaEvolve and the Darwin Gödel Machine.The limits matter as much as the gains. Jiang explains why the experiment did not establish that the system had become a better improver. The conversation closes with open-ended search, human-designed primitives and Parameter Golf: where does the next useful idea come from when the agent is searching inside a space that people designed?---TIMESTAMPS:00:00:00 Eight days of self-improvement: what counts?00:03:25 AIDE and the puzzle of useful spaghetti code00:08:38 Four levels of recursive self-improvement00:12:02 What AIDE 85 changed and how it was tested00:20:04 AlphaEvolve, Darwin Gödel Machine and the RSI claim00:26:21 Reward hacking and the limits of detection00:33:09 Open-ended search, harness tuning and creativity00:39:43 Parameter Golf and the limits of self-improvement---REFERENCES:organization:[00:00:30] Weco AIhttps://www.weco.ai/other:[00:00:33] AIDE²: The First Evidence of Recursive Self-Improvementhttps://www.weco.ai/blog/first-evidence-of-recursive-self-improvement[00:14:11] Faulty reward functions in the wildhttps://openai.com/index/faulty-reward-functions/[00:29:59] The Hugging Face incident and the road aheadhttps://openai.com/index/hugging-face-incident-and-the-road-ahead/tool:[00:03:29] AIDEhttps://github.com/WecoAI/aideml[00:04:29] MLE-benchhttps://github.com/openai/mle-bench[00:04:33] ALE-Benchhttps://github.com/SakanaAI/ALE-Bench[00:04:52] WeatherBench 2https://github.com/google-research/weatherbench2[00:08:18] ReActhttps://react-lm.github.io/[00:39:43] Parameter Golfhttps://github.com/openai/parameter-golfpaper:[00:20:08] AlphaEvolve: A coding agent for scientific and algorithmic discoveryhttps://arxiv.org/abs/2506.13131v1[00:21:35] Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agentshttps://arxiv.org/abs/2505.22954v3[00:23:45] Hyperagentshttps://arxiv.org/abs/2603.19461v1[00:27:01] SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agentshttps://arxiv.org/abs/2605.21384book:[00:33:14] Why Greatness Cannot Be Planned: The Myth of the Objectivehttps://link.springer.com/book/10.1007/978-3-319-15524-1---LINKS:https://app.rescript.info/share/3a9dc6189cb539c6a05fcc4f75c101b3PDF:https://app.rescript.info/api/public/sessions/9eda60ede2b31c92/pdf
- Frank Hutter, co-founder of Prior Labs, talks about TabPFN, a tabular foundation model that makes predictions in a single forward pass, and the research behind it.
TabPFN is pre-trained on synthetic datasets drawn from a prior over structural causal models, rather than on real data. At prediction time it takes the whole training table as context and outputs an approximation of the Bayesian posterior predictive distribution, without per-dataset training or hyperparameter search. Frank explains how this grew out of his earlier work on AutoML and neural architecture search, how the priors are built and revised, and why tabular data was hard for deep learning for so long.
The conversation also covers the TabArena benchmark, how the architecture changed from TabPFN v1 to v3, scaling to larger tables, using the model with coding agents, test-time compute, Google's TabFM, causal inference and interventions, and relational data. At the end, a short update Frank recorded after the interview covers the TabPFN-3.5 release.
Prior Labs:
TabPFN-3.5: https://priorlabs.ai/tabpfn-3-5
https://priorlabs.ai/careers
TOC:
00:00 Introduction
00:44 Welcome and Frank's background
02:05 Why tabular data was hard for deep learning
10:17 Pre-training on synthetic data
12:52 The TabArena benchmark
19:28 From AutoML to neural architecture search
26:34 TabPFN as a learned algorithm
30:50 Bayesian prediction in one forward pass
39:37 Scaling to larger tables
47:48 Using TabPFN with coding agents
57:47 Output heads and architecture from v1 to v3
1:05:29 Test-time compute and adaptation
1:13:32 Google's TabFM
1:16:53 How the priors are designed
1:18:40 Correlation, causation and interventions
1:35:22 Relational and multimodal data
1:38:31 Use in organisations
1:46:38 The open research arm
1:50:21 Update: TabPFN-3.5
REFS:
TabPFN v2, Nature (Hollmann et al., 2025)
https://www.nature.com/articles/s41586-024-08328-6
Transformers Can Do Bayesian Inference (Müller et al.)
https://arxiv.org/abs/2112.10510
TabArena (Erickson et al.)
https://arxiv.org/abs/2506.16791
AutoGluon-Tabular (Erickson et al.)
https://arxiv.org/abs/2003.06505
Beyond IID: How General Are Tabular Foundation Models, Really?
https://arxiv.org/abs/2606.30410
Neural Architecture Search: A Survey (Elsken, Metzen & Hutter)
https://arxiv.org/abs/1808.05377
Auto-WEKA (Thornton et al.)
https://www.cs.ubc.ca/~hutter/papers/AutoWEKA-KDD2013.pdf
TabPFN v1 (Hollmann et al., 2022)
https://arxiv.org/abs/2207.01848
TabPFN-3 technical report
https://arxiv.org/abs/2605.13986
TabPFN-2.5 report
https://arxiv.org/abs/2511.08667
CAAFE (Hollmann et al.)
https://arxiv.org/abs/2305.03403
TabICL (Qu et al.)
https://arxiv.org/abs/2502.05564
TabICLv2 (Qu et al.)
https://arxiv.org/abs/2602.11139
Google TabFM
https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/
TALENT benchmark (Ye et al.)
https://arxiv.org/abs/2407.00956
Do-PFN (Robertson et al.)
https://arxiv.org/abs/2506.06039
CausalPFN (Balazadeh et al.)
https://arxiv.org/abs/2506.07918
Causal Foundation Models with Partial Graphs (Reuter et al.)
https://arxiv.org/abs/2602.14972
RelBench (Robinson et al.)
https://arxiv.org/abs/2407.20060
RelArena-α, TabPFN-Rel and RPI
https://arxiv.org/abs/2608.16319
TabPFN on GitHub
https://github.com/PriorLabs/TabPFN
TabPFN-3.5 technical report
https://arxiv.org/abs/2609.17895
Otto Group Product Classification Challenge (Kaggle, 2015)
https://www.kaggle.com/competitions/otto-group-product-classification-challenge
---RESCRIPT:https://app.rescript.info/share/e99676c25ee6189fbf54c9be07eb623e
Meer Technologie podcasts
Trending Technologie -podcasts
Over Machine Learning Street Talk (MLST)
Welcome! We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis. Our approach is unrivalled in terms of scope and rigour – we believe in intellectual diversity in AI, and we touch on all of the main ideas in the field with the hype surgically removed. MLST is run by Tim Scarfe, Ph.D (https://www.linkedin.com/in/ecsquizor/) and features regular appearances from MIT Doctor of Philosophy Keith Duggar (https://www.linkedin.com/in/dr-keith-duggar/).
Podcast websiteLuister naar Machine Learning Street Talk (MLST), AI Report en vele andere podcasts van over de hele wereld met de radio.net-app

Ontvang de gratis radio.net app
- Zenders en podcasts om te bookmarken
- Streamen via Wi-Fi of Bluetooth
- Ondersteunt Carplay & Android Auto
- Veel andere app-functies
Ontvang de gratis radio.net app
- Zenders en podcasts om te bookmarken
- Streamen via Wi-Fi of Bluetooth
- Ondersteunt Carplay & Android Auto
- Veel andere app-functies


Machine Learning Street Talk (MLST)
Scan de code,
download de app,
luisteren.
download de app,
luisteren.




























