148 afleveringen
SkyDiscover with Alexander Krentsel, Shu Liu, Shubham Agarwal, and Mert Cemri - Weaviate Podcast #147!
05-10-2026 | 51 Min.Shu Liu, Alexander Krentsel, Mert Cemri, and Shubham Agarwal, PhD students at U.C. Berkeley's Sky Computing Lab, join the Weaviate Podcast to discuss SkyDiscover and what it means for AI to build entire systems, not just code! AI coding went from Copilot-style line completion to function- and file-level generation with Claude Code. The next step is to optimize or synthesize systems that used to take PhD-level expertise, such as databases, inference engines, and key-value stores. With coding agents, "just-in-time systems" become possible: software tailored to one workload, free of the "generality tax" of general-purpose systems.The conversation centers on synthesizing a key-value store, where workloads vary in read/write ratio, hot keys, consistency, reliability, and latency versus throughput goals. Running on curated and production traces from Tencent and Twitter, the agents designed new caching and hierarchical memory schemes that reached up to 4x better performance than specialized systems from Microsoft. The hard part is trust. On YCSB traces, early builds showed 30–40x throughput because Claude was reward hacking: it worked out values from the keys. SkySynth acts as a "vaccine" against this. It asks the user questions to build the spec, then takes either a test-driven route or a formal verification route. In the test-driven route, a test only counts if it can tell a correct reference implementation from an incorrect one.From there, the discussion covers how the spec decides the programming language (Rust for memory safety, C for raw performance, Lean, Rocq, CompCert, and Verus for proofs). A surprising result: building on top of Microsoft FASTER gave lower throughput than synthesizing from scratch. The panel's advice is to "skate to where the puck is going" and let capable models design without constraints. The panel also imagines hosted systems resynthesized daily on recent traffic, and PRDs as the spec for consumer products.The episode closes on optimization. Unlike text optimization with GEPA, 80–90% of system optimization time goes to evaluation, and vLLM alone has more than 25,000 tests. That makes finding the right problem, building proxy evaluators, and generating adversarial workloads a meta-problem in its own right.
Chapters
0:00 Welcome Shu, Shubham, Alexander, and Mert!
2:22 An Overview of SkyDiscover
7:55 Specialized Database Systems
12:33 DSPy for Systems
23:23 Formal Verification of AI-Generated Systems
26:34 Programming Languages and Vibe Coding
32:48 Building Systems from Scratch
40:26 Software for Everyone!
45:26 GEPA for Vibe Coding- Siddharth Gollapudi, a researcher at UC Berkeley, joins the Weaviate Podcast to discuss in-context retrieval and his paper "Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale." The idea is simple and radical: instead of encoding documents into embeddings and building a nearest neighbor index, put the entire corpus into an LLM's context window and use attention itself as the retriever. The conversation opens by separating this from generative retrieval, which memorizes documents into model parameters and needs gradient updates every time the corpus changes, and from listwise re-ranking, which Siddharth frames as an easier subset of first-stage retrieval. Lost in the middle, he argues, turned out to be much less of a problem with today's long-context models.The discussion then dives into BlockSearch, a 0.6B parameter in-context retriever trained on MS MARCO (cleaned with RLHN) plus a mix of BEIR datasets. Reaching 500 documents was relatively easy; scaling to 2,500, 5,000, and 10,000 documents, or roughly a million tokens, is where things break. Siddharth explains why random document codes beat positional codes (they force the model to actually read the documents), and how an on-policy loss that corrects the model's own rollouts keeps it disciplined when hard negatives pull it off course midway through generating a code.From there, the conversation moves to attention dilution: as the corpus grows, the softmax denominator swamps the relevant document's score. Two fixes help, length-dependent temperature scaling and sparse attention that drops irrelevant documents before attention runs, bringing BlockSearch roughly level with dense retrieval at a million tokens. This opens a big question for vector databases: is there a sublinear, ANN-style version of attention that can be trained end to end, in the spirit of ReFrag and ColBERT's late interaction?The episode closes on timelines for LLM-based re-rankers, the latency trade-offs of high-latency search, "No More Free Lunch" and whether long-context LLMs can eat the database, recursive language models, and Siddharth's excitement for unsupervised notions of relevance that could help models make genuinely new discoveries, such as proving open theorems.
Chapters0:00 Welcome Siddharth!1:21 In-Context Retrieval8:16 Drowning in Documents and Reranker Scaling13:05 BlockSearch, 0.6B In-Context Retriever28:21 Vector Databases for LLM Inference40:46 Timelines for In-Context Retrievers46:54 Will LLMs eat Databases?51:38 Exciting Directions for AI humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145!
21-09-2026 | 51 Min.Alexis, Manya, and Niloofar from humans& join the Weaviate Podcast to introduce Persimmon, a user model built to simulate how humans actually behave in multi-turn, multi-party conversations. Persimmon is explicitly not an assistant, a companion, or a Character AI-style stand-in, it is a research preview aimed at faithfully capturing the distribution of human behavior. This includes the natural friction of frustration, excitement, and group dynamics that assistant chatbots trained to be helpful never exhibit. Alexis, Manya, and Niloofar bring a striking mix of backgrounds to the problem: AI tutoring and student modeling, programming languages for high-performance computing, and privacy and information-flow research at Carnegie Mellon. The conversation opens with whether the Turing test is solved. Humans& runs a distributionally grounded, multi-turn version where the judge sees many examples of human and AI behavior. Frontier models fool it less than 5% of the time, while Persimmon reaches roughly 20% against a 50% ceiling. From there, the discussion dives into training for non-verifiable tasks: why rubrics-as-rewards approaches invite reward hacking, why the team refuses to impose its own theory of human behavior, and how distribution matching, with the multi-turn Turing test as a North Star metric operating in an implicit feature space, rather than Earth mover's distance over hand-picked features anchors both training and evaluation. They walk through evaluating on real human interaction data like the TIDES meeting transcripts and the TutorMoments tutoring dataset, and why role-played or scripted dialogue doesn't count.The discussion then moves into theory of mind and world models as twin goals, with Persimmon enabling multi-agent environments where assistants get realistic human feedback at training time. The podcast further covers the choice of NVIDIA's Nemotron 3 Ultra and why starting from a base model matters: post-training causes mode collapse, you can't prompt-optimize your way out of it, and injected randomness drifts away over long rollouts. The conversation lands on what excites each guest next: personalized tutors, models that balance overlapping human goals, and training paradigms with long-term social pressures.
Chapters
0:00 Welcome Niloofar, Manya, and Alexis!2:22 An Overview of Persimmon6:08 Solving the Turing Test13:36 User Models and AGI16:00 RL with Non-Verifiable Rewards20:35 Distribution Matching28:00 Collecting Human Data32:32 Theory of Mind in AI37:40 NVIDIA Nemotron 3 Ultra43:49 Prompt Optimization46:00 Exciting Directions for AI- Alex Ledbetter is the founder of Delegance Brokerage, an AI-native commercial insurance brokerage he built entirely on his own. The conversation opens with his origin story: an internship on a political risk, credit, and bond underwriting team in New York, where he watched a three-trillion-dollar portfolio run off the back of Excel and realized commercial insurance brokers earning 15–35% annual commissions could be disintermediated the way Robinhood disintermediated retail stock brokers. After raising $200K and securing 190 state licenses in 45 days, he spent seven months teaching himself to build with AI coding tools. Alex compares this process to hammering sheet metal, walking through his own product flow 700 times before showing it to a client.From there, the discussion dives into the architecture. Clients drop in insurance binders approaching a thousand pages, and a document processing pipeline classifies up to 50 commercial insurance document types, vectorizes everything into Weaviate, builds a table of contents per document for multi-hop traversal search, and runs trailing async extractors that pull structured fields into Postgres, processing 80 pages per minute per concurrent upload by converting PDFs into images for vision APIs rather than relying on flat OCR. Three million carrier appetite rules then route each client to the right insurer. Since only 7 of his 27 carrier partners have APIs, browser agents log into carrier portals, handle one-time passwords, and fill out underwriting questionnaires by querying Weaviate and Supabase in real time and monthly compliance agents renew licenses across 50 states.The conversation moves into agent harnesses: multi-model developer pipelines using Codex, Claude, and BugBot that have opened over 900 PRs a month, an email operating system with a chief of staff over iMessage, and how Weaviate unifies memory across web, iOS, Slack, email, and phone. It lands on what's next: a fresh fundraise, SOC 2, and running lean with embedded partnerships instead of brokers.
Chapters:0:00 Welcome Alex!0:38 Founding Story of Delegance Brokerage
6:38 Insurance Brokers and AI
12:15 A Knowledge Base for Insurance
24:20 Vibe Coding
28:50 Browser Agents
34:14 Email Agents
37:30 What is an Agent Harness?
49:00 Weaviate and Postgres Database Design
55:15 The Future of Delegance Brokerage - Sam O'Nuallain joins the Weaviate Podcast to discuss AutoIndex, research from UMass Amherst and Databricks on learning representation programs for retrieval. Instead of tuning the retriever or re-ranker, AutoIndex asks how the data itself should be represented: a two-agent system, an analysis agent and a code agent, writes and refines Python programs that chunk, enrich, and reorganize a corpus to optimize downstream metrics like Recall and nDCG.The conversation opens with why indexing is such a natural target for code optimization. Frontier LLMs are exceptional at writing code, a representation program applies cheaply across an entire corpus without passing every document through an LLM, and code is verifiable. Every hypothesis the code agent proposes is gated against a validation set before it is accepted. From there, the discussion dives into the optimization signal. The analysis agent uses tools to read documents, query the retriever, and inspect where gold documents rank, turning a bare score like "recall went up" into rich natural language feedback about why a representation is failing, echoing ideas like GEPA's reflective metrics and the value of small-margin positives for training re-rankers.That leads into why BM25 pairs so well with agents: its lexical transparency makes failures easy to diagnose, illustrated by case studies from the CRUMB benchmark. This includes LaTeX formatting errors sinking Stack Overflow retrieval and Tip-of-the-Tongue movie search, where AutoIndex learned to repeat plots to up-weight terms and expand documents with synonym dictionaries. The conversation moves through connections to document enrichment methods like Anthropic's contextual retrieval, doc2query, and EnrichIndex, generalizing representation programs to text-to-SQL schemas and data lakehouses. The podcast concludes by discussing Sam's lessons transitioning from research to production AI engineering: loop engineering, QA, and evals. It lands on the directions that excite Sam most: harness design, continual learning, and memory as a retrieval problem, squeezing more out of the models we already have without touching the weights.
Chapters
0:24 An Overview of AutoIndex
4:04 Retrieval Indexing as Code Optimization
8:47 Optimizing Chunking and Database Schemas
15:57 Feedback for Search Optimization
23:57 Future Directions for AutoIndex
27:34 Document Enrichment for RAG
36:32 Web Search vs. Databases
40:40 AI Engineering
49:30 Exciting Directions for AI
Meer Technologie podcasts
Trending Technologie -podcasts
Over Weaviate Podcast
Join Connor Shorten as he interviews machine learning experts and explores Weaviate use cases from users and customers.
Podcast websiteLuister naar Weaviate Podcast, Acquired en vele andere podcasts van over de hele wereld met de radio.net-app

Ontvang de gratis radio.net app
- Zenders en podcasts om te bookmarken
- Streamen via Wi-Fi of Bluetooth
- Ondersteunt Carplay & Android Auto
- Veel andere app-functies
Ontvang de gratis radio.net app
- Zenders en podcasts om te bookmarken
- Streamen via Wi-Fi of Bluetooth
- Ondersteunt Carplay & Android Auto
- Veel andere app-functies


Weaviate Podcast
Scan de code,
download de app,
luisteren.
download de app,
luisteren.































