Podcasts about Llama

Species of wooly domesticated mammal

  • 3,833PODCASTS
  • 9,603EPISODES
  • 38mAVG DURATION
  • 2DAILY NEW EPISODES
  • Aug 30, 2026LATEST
Llama

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about Llama

Show all podcasts related to llama

Latest podcast episodes about Llama

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE 28 de agosto de 2026

Jay Fonseca

Play Episode Listen Later Aug 28, 2026 24:18


PODCAST LAS NOTICIAS CON CALLE 28 de agosto de 2026 - EEUU negocia quedarse con una tajada "masiva" del petróleo de Venezuela, Venezuela negocia quedarse con 90 mil millones de barriles de petróleo y que Venezuela salga de la OPEP Ayuda federal tardaría hasta un año por sequía y desastre federal - El Vocero Crisis de agua inunda la línea PAS - El Vocero Un momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequía. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado. Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#windmarhome #windmarhome #incluyeauspicioCorrección admite que no hay razón para cancelar contrato ordenado por la gobernadora cancelar y debe 112 millones a empleados en retroactivo - El Nuevo Día Tormenta Dolly sigue camino hacia acá para lunes - NHCSolX cesantea ~200 en Aguadilla tras suplidos no poder entregar materiales - Jay FEmbalses al borde: Carraízo en 37.27m, a 26 cm del nivel de "control" (37.00m); La Plata bajó 11 cm a 45.05m; 14 de 19 embalses reportaron bajas. 180,000 abonados en racionamiento de 48 horas - El Vocero Yanira Raíces incumplía requisitos desde 1999, por eso pide distinción especial - El Nuevo Día Jefe de DACO plantea correr para alcalde de Cataño y tiene comité para político mientras es secretario - El Nuevo DíaRegresa SuperPAC Democracia es Prosperidad - El Nuevo Día Norma Burgos pide break al Senado para entregar datos de los últimos días de Domenech y García Posible cirugía Tommy John para Fernando Cruz Boston Scientific sufre ciberataque global, afectó su capacidad de procesar y enviar pedidos; reportado a la SECFederación agrícola en quiebra - El Nuevo Día Trump renombra Lago Ontario "Lake America":7 de 14 unidades de respuesta rápida fuera de servicio, se las compramos a los dueños de Genera y no sirven, Negociado dice que va a investigar - El Nuevo Día Escala la pelea con Canadá, Trump impuso aranceles de 50% a una batería de productos canadienses; Canadá respondió con 50% al alambre de cobre, madera y carbón vegetal gringos, y ofreció C$504M para atraer a 64 científicos estadounidenses - Bloomberg Hoy el jefe de la FED da su primer discurso Jackson Hole a las 10 Nvidia, imparable. $60 billones de ganancia trimestral, esto es algo nunca antes visto Lilly logra que aprueben Mounjaro para condiciones cardíacas - Bloomberg Irán cumple 6 meses de guerra. Hormuz "reabre" y el crudo se estabiliza, pero el diésel en EEUU sigue a $5.65 el galón250+ negocios piden alivio de hasta $10,000 a Turismo, 254 solicitudes de las 450-500 estimadas - El Nuevo Día Trump vs. Canadá. Aranceles de 50%, Canadá responde, y Trump renombra el Lago Ontario "Lake AmericaLOS DATOS DEL DÍABrent: $87.65/barril  Diésel EEUU (retail): $5.65/galón (+19.8¢ en la semana)  ·  S&P 500: 7,730.99 (+0.7%)  ·  Dow: 53,569.44 (+0.2%)  ·  Nasdaq: 26,541.35 (+1.6%)Bono 10 años: 4.67%  ·  Euro/USD: 1.1656  ·  Gas natural (Henry Hub): ~$3.00/MMBtu (Europa en máximos de 3 años)  ·  Hipoteca 30 años: 6.66%  ·  Oro: mineras +43% en agosto

Homilias – Casa para tu Fe Católica
LA GRACIA 2026/08/29 Juan Bautista nos llama a preparar el corazón para Cristo

Homilias – Casa para tu Fe Católica

Play Episode Listen Later Aug 27, 2026


MARTIRIO DE SAN JUAN BAUTISTA Juan predicaba y bautizaba para anunciar que el Mesías estaba cerca y llamar a todos a la conversión. Su vida nos recuerda que necesitamos un corazón firme, que no se someta al pecado. http://traffic.libsyn.com/fraynelson/sjbm021a.mp3

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE 25 de agosto de 2026

Jay Fonseca

Play Episode Listen Later Aug 25, 2026 19:46


PODCAST LAS NOTICIAS CON CALLE 25 de agosto de 2026 - Rivera Schatz se abraza y besa con cliente de Politank, les ayuda a cuadrar el caso ante los tribunales para quedarse guisando en Corrección Guerra de billetudos, dueño de hotel de mega ricos exige demoler hotel de clase media en construcción en Dorado, millones de cooperativas en riesgo - El Nuevo Día El día de los silencios: Power Expectations no habla, la AAA no entrega datos y nadie sabe cuánto cuesta la pensión de $1,000 que proponen a aumentar El Senado, por instrucción de Rivera Schatz, acudirá al tribunal para obligar a la AAA a entregar información requerida desde el 17 de junio en Comisión Total Guisando Carlitos Bermúdez según el CPI Siempre innovando y con los mejores beneficios, MCS Personal Directo te ofrece cubiertas accesibles para que cuides de tu salud y la de los tuyos.Con una amplia red de proveedores de más de 15,000 médicos de libre selección. Reembolso de hasta $40 mensuales por membresía a un gimnasio o por un entrenador personal debidamente certificado. Asistencia en el hogar para servicios de cerrajería, plomería y electricidad de hasta $350 por evento hasta 4 veces al año.¡Únete HOY a la gran familia de MCS!¡Salud que completa tu vida! Llama al 787.945.1259 y oriéntate.Endoso pagado#MCS#incluyeauspicio Jgo radicó como proyecto (PS 1428 / PC 1399) para subir la pensión mínima de los retirados del gobierno a $1,000, no sabe de dónde van a salir los fondos - El Vocero Diésel en récord: ~$5.45/galón, casi $100 sobre el crudoABRE PR reporta que 32 de 74 municipios (43%) gastaron más de lo que ingresaron en 2024Cita mañana miércoles al director de la AAPP, Josué Colón Ortiz para que explique el contrato de Power Expectations y 3PPO en la Cámara - El Vocero Aprobado por Cámara y Senado el PC 911 que prohíbe el “voceteo” nocturno y en zonas escolares o de tranquilidad - El Vocero Senado aprobó crear la Junta Examinadora de Servicios Estéticos - El Vocero Disputa por la construcción del Hilton Garden Inn de Dorado   Boricuas enamorados de Colombia, 60,000 viajeros de PR visitaron Colombia en 2025, casi +60% y van 2-3 veces - El Nuevo Día Le sueltan rienda a Yovngchimi, federales dejan en manos de oficial probatorio - El Vocero 1 de cada 4 jugadores de la NFL tienen daño cerebral - NYTTrump podría limitar el voto por correo, permite la Corte Suprema - NYTUSA no sancionó a China todavía, quedaron en hacerlo - Bloomberg La controversia de las cámaras Flock se dispara en todo USA - WSJTrump va a revocar 200 mil visas, lo más realizado en la historia - Axios Más café se pierde por falta de ano de obra que por seuquía  LOS DATOS DEL DÍA (cierre lunes 24 ago)Brent~$93/bbl (▼ ~1.3%)WTI$82.80 (▼ 2.6%)Diésel retail EEUU~$5.45/gal (récord)Gasolina retail EEUU~$4.43/galS&P 5007,652.86 (▼ 0.28%)Dow Jones53,417.16 (▲ 0.26%)Bono 10 años4.70%Euro / USD1.1664Gas natural (Henry Hub)~$2.81/MMBtuHipoteca 30 años6.65%Bitcoin>$80,000

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE 20 de agosto de 2026

Jay Fonseca

Play Episode Listen Later Aug 20, 2026 25:01


PODCAST LAS NOTICIAS CON CALLE 20 de agosto de 2026 - FBI investiga transacción de Power Expectations - El Vocero Gobiernos extranjeros ya no quieren comprar tanto la deuda de USA, Tesoro compra bonos federales  - Axios Hoy 11:00 AM Directorio PNPFederales cogen a Agricultura con fallas - OIG halla 21 deficiencias "sistémicas" en Seguros Agrícolas — expedientes incompletos, contratos sin cláusulas, pagos sin evidencia y solo buscaron 201 expedientes de miles - El Vocero La Década Perdida": la Junta a 10 años de PROMESA y dicen que la Junta no ha logrado resolver los problemas de PR - El Vocero Un momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequía. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado. Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#windmar#incluyeauspicio $772,400 en un vagón: 22 meses estacionado en Bayamón mientras Vieques sigue sin hospital y equipo fue comprado para dejarlo allí - El Nuevo Día Politank sin pagar las multas del DDEC y los 80 mil todavía - El Nuevo Día PIP propone que si Ignoras una orden del DACO te afecte el crédito - El Nuevo Día JGo le recuerda el pasado a Thoommy de cuando él era presidente - El Nuevo Día Trump anuncia guerra económica contra cualquier país que ayude a Irán - CNBC Una interrupcion electrica saco cuatro bombas en la represa La Plata FEI entrega informe sobre Jessika Padilla expresidenta de la CEE y básicamente aquí aparenta ser que se usó el poder para perseguirla por el PNP - El Nuevo Día Hoy PNP radica proyecto para la entidad - El Nuevo Día Guerra economica contra Iran: Emiratos corta lazos y sube el crudo • • Moderna se duplica en bolsa por vacuna contra el melanoma

Jorge Ramos Y Su Banda
Barcelona ya tiene su ‘9' y se llama Lamine Yamal

Jorge Ramos Y Su Banda

Play Episode Listen Later Aug 14, 2026 87:53


En Ahora o Nunca, la mesa discute si el Barcelona ya está lo suficientemente armado para ser considerado candidato indiscutible a ganar LaLiga sin un '9' nominal, al señalar que Lamine Yamal puede cumplir, y quizá lo haga, con esa función, pero no descartan la posibilidad de que el club culé sea considerado el mejor de Europa si consigue fichar a un delantero nato. Learn more about your ad choices. Visit podcastchoices.com/adchoices

Así las cosas
Kenia López Rabadán llama a frenar agresiones en San Lázaro tras choque entre diputados de PRI y Morena

Así las cosas

Play Episode Listen Later Aug 13, 2026 6:43


La presidenta de la Mesa Directiva de la Cámara de Diputados pide civilidad legislativa tras el altercado entre Carlos Eduardo Gutiérrez Mancilla y Arturo Ávila, y advierte que el Congreso debe mantener el debate en el terreno de las ideas.

Así las cosas
Alejandro Moreno defiende al PRI ante investigaciones, choques legislativos y llama a una oposición unida

Así las cosas

Play Episode Listen Later Aug 13, 2026 16:41


El dirigente nacional del PRI habla con Gabriela Warkentin sobre las confrontaciones en el Congreso, las investigaciones contra excolaboradores en Campeche y la posibilidad de construir coaliciones opositoras rumbo a los próximos procesos electorales.

Atareao con Linux
ATA 822 PowerPoint HA MUERTO! Genera presentaciones con IA en 15 segundos

Atareao con Linux

Play Episode Listen Later Aug 13, 2026 28:23


Hace unos meses empecé a usar presentaciones para grabar el podcast, y enseguida me di cuenta de que el verdadero problema no es pensar el contenido, sino maquetarlo. Pasaba más tiempo ajustando fuentes, colores y transiciones que preparando lo que realmente quería contar. Así que me puse a buscar una solución, y lo que encontré me ha cambiado el flujo de trabajo por completo.En este episodio te cuento cómo he montado typst-ia, un script en Python que genera presentaciones completas en segundos. Le dices un tema, la inteligencia artificial se encarga del contenido, y Typst lo convierte en un PDF impecable. Todo desde la terminal, sin abrir PowerPoint ni Google Slides, sin suscripciones mensuales, y con un control total sobre el resultado.Typst es un sistema de composición moderno escrito en Rust que compila en milisegundos. Sí, has leído bien, milisegundos. Comparado con LaTeX Beamer, que tarda 5 o 10 segundos en compilar, Typst es un antes y un después. Además, su sintaxis es mucho más limpia y fácil de aprender. En el episodio lo comparo con LaTeX y con Markdown, y te cuento por qué creo que Typst se está convirtiendo en el estándar para presentaciones técnicas.La clave del proceso está en el system prompt. Incrusto el template real de la presentación dentro del prompt que le envío a OpenRouter, y la IA genera código Typst válido sin necesidad de retoques. Uso DeepSeek Chat por defecto —cuesta unos 14 céntimos por millón de tokens de entrada, que vienen a ser cientos de presentaciones por menos de un euro—, pero también puedes usar Claude Sonnet, Gemini Flash o Llama 3.3 si necesitas más calidad o prefieres un modelo concreto.El script completo son unas 200 líneas de Python sin frameworks, solo con la librería requests. Te explico paso a paso cómo funciona el pipeline: lee el template, construye el prompt, llama a OpenRouter, limpia la respuesta, escribe el archivo .typ, lo compila a PDF y lo abre en el visor. Y todo con flags para personalizar el número de diapositivas, el modelo, el nombre del archivo y hasta los reintentos si la compilación falla.Para rematar, hago una demo en vivo generando una presentación desde cero. Ves cómo en cuestión de segundos pasamos de una idea a un PDF listo para proyectar. Y lo mejor es que el resultado es texto plano, versionable con Git, editable con cualquier editor, y sin ningún tipo de lock-in. Si mañana quieres cambiar algo, abres el .typ y lo tocas.Si eres de los que hacen presentaciones técnicas, charlas, workshops, o simplemente quieres automatizar una tarea tediosa, este episodio te va a gustar. Y si nunca has oído hablar de Typst, te vas a llevar una sorpresa.Capítulos del episodio:0:00 - Introducción: presentaciones con Typst e IA2:52 - El problema de las presentaciones tradicionales5:20 - Typst: el sistema de composición moderno7:34 - Typst vs LaTeX vs Markdown8:31 - Instalación de Typst9:28 - Plantillas para presentaciones con Typst12:52 - OpenRouter y el prompt para la IA15:28 - El script Python: el pipeline completo17:41 - Demo en vivo: generando una presentación24:17 - Conclusiones y despedida

ABC Noticias
Exgobernador de Guerrero busca evitar vinculación por caso Ayotzinapa

ABC Noticias

Play Episode Listen Later Aug 12, 2026 12:29


En mas informacion: Llama titular de Sader al sector industrial a convertirse en socios estratégicosMéxico supera las 100 mil personas en prisión preventiva: 4 de cada 10 internos sin sentenciaGuanajuato baja 76% homicidios, pero sigue liderando violencia letal en MéxicoTortilla sube de precio en Tamaulipas: el kilo llega hasta 32 pesosSecretaría de Salud descarta casos de diarrea explosiva en ZacatecasSLP lidera reducción de homicidios dolosos en México con caída de 80.5%Shein apunta a debutar en la bolsa de Hong Kong este miércoles¡Vergüenza tras vergüenza! Pumas vuelve a hacer el ridículo luego de perder el punto extra ante Columbus“La Granja VIP” 2026: quiénes son los participantes del reality Hosted on Acast. See acast.com/privacy for more information.

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE 11 de agosto de 2026

Jay Fonseca

Play Episode Listen Later Aug 11, 2026 18:35


PODCAST LAS NOTICIAS CON CALLE 11 de agosto de 2026 - Países del Golfo Pérsico admiten que Irán es quien manda en Hormuz, Trump dice que manda él - WSJMega tapón de Caguas a San Juan por camión que se trepó en valla y deja combustible allí - WAPA Hacen falta 15 pulgadas de lluvia para resolver el problema de Carraízo - El Nuevo Día Reunión de hoy de Norma Burgos y Rivera Schatz hoy a las 3:30 - El Vocero Genera quiere hacer plantas desalinizadoras para enfriar plantas de energía - El Vocero Trump vuelve a dar break de las leyes de cabotaje par apoder mover Gas Natural dentro de USA - El Vocero Esperan lluvias y tronadas para el centro y oeste por débil onda tropical sobre PR - Metro Montones de demandas contra redes sociales por adicción siguen su curso tras decisión de Circuito de Apelaciones - Reuters Trump le exige "reparaciones" a Irán y sube la tensiónUn momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequía. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado. Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#windmarhome#incluyeauspicio Gobernadora anuncia que anunciará algo importante con la AAA - X JGo pide a Rivera Schatz pasar la página - X Admiten que el ERP no estaba listo porque no han adiestrado a empleados públicos para su implementación - El Vocero El mes más caliente en la historia reportada ha sido julio, peor que el famoso dust bowl de 1936 - AP A 38 billetes el galón de gasolina en Cuba (en PR a 4) - Axios Trump cambia el régimen de vacunas y plantea que se vacune menos a niños y dividir algunas vacunas - Axios Roban en casa y matan a tiros a la Pitbull que estaba en el cuarto principal - Primera HoraSe cumplen un año de muerte de Gabriela Nicole - Primera Hora Syria sentencia a muerte a Bashar al Assad - Reuters Ya van 160 muertos en Colombia por terremoto de 7.4 - Reuters Sigue en espera el W de Vieques de demanda en tribunal - El Nuevo Día LOS DATOS DEL DÍA (cierre lunes 10 de agosto) Brent$87.69  ▲ +4.95% Diésel PR (detal)~$1.18–$1.27/litro S&P 5007,753.11  ▼ -0.06% Dow Jones53,975.98  ▼ -0.11% Bono 10 años4.666% Euro/USD1.1547  ▼ -0.10% Gas natural (Henry Hub)~$2.66/MMBtu Hipoteca 30 años6.823%

Templo Mayor
TEMPLO MAYOR : Miopía selectiva le llaman

Templo Mayor

Play Episode Listen Later Aug 11, 2026 3:13


Llama la atención el silencio de Morena ante revelación de que un precandidato tiene en su equipo a un señalado operador de delincuencia.

Relatos del lado oscuro
Cuando te llama un fantasma Relatos del lado oscuro

Relatos del lado oscuro

Play Episode Listen Later Aug 10, 2026 61:52 Transcription Available


Suena el teléfono isistentemente... pero la voz que escuchas no es un cliente sino tu padre que ha muerto tiempo atrás....qué haces.Conviértete en un supporter de este podcast: https://www.spreaker.com/podcast/relatos-del-lado-oscuro--5421502/support.#relatos de misterio #relatos de terror #historias de miedo #asesinos #terror parapsicológio #joseramoncantalapiedra

GotTechED
10 Tools for Specialized AI and Productivity

GotTechED

Play Episode Listen Later Aug 10, 2026 36:03


Edtech Throwdown Episode 221: 10 Tools for Specialized AI and ProductivityWelcome to the EdTech Throwdown. This is Episode 221 called 10 Tools for Specialized AI and Productivity. In this episode we'll be talking edtech tools as we bring you a list of some atypical AI platforms that might not otherwise come up for educators. This is another episode you don't want to miss. Check it out.Segment 1:Happy August, happy end of summer breakWe're almost fully re-charged and starting to think about things like productivity againSegment 2:Nick: abacus.ai: Abacus.AI is the world's first AI super-assistant tailored for enterprises and professionals. We offer two products: ChatLLM for professionals and small teams and Abacus.AI Enterprise for enterprises and companies. ChatLLM is a multi-modal, multi-device super-assistant that can handle many tasks and completely transform your life. You can access all of the SOTA LLMs, analyze documents, do data analysis, generate code, search the web, create images, and much more. It's your all-in-one AI assistant and increases individual productivity by 15% to 75%. Abacus.AI Enterprise is a state-of-the-art generative AI platform that combines the AI super assistant available to all your employees with an AI brain that can connect to your enterprise software systems, automate business processes, increase revenue, and be a powerful force multiplier. Our AI engineer can build AI Workflows and chatbots to automate critical processes. Our Enterprise product comes with single sign-on, multiple deployment options and checks all the security and compliance boxes.Guise:studley.aiWhy:Specialized AI platforms designed for more complex, data-heavy, or professional-grade automation and modeling.Nick: openalternative.co: What is the difference between open source and proprietary? Proprietary AI Tools: These are "closed-source" models developed and owned by specific companies (e.g., OpenAI's GPT-4/5, Google's Gemini, Anthropic's Claude). You typically access them via an API or a web interface. The underlying code, training data, and exact model weights are a "black box" hidden from the public. Open-Source (Open-Weight) AI Tools: These are models whose code and weights are publicly shared (e.g., Meta's Llama series, Mistral, Qwen, and DeepSeek). Anyone can download them, look under the hood, modify them, and run them on their own hardware or private cloud. Guise: The Most Comprehensive List of FREE Online Tools for Teachers Why:Both promote open-source ethics, whether finding alternatives to paid software or using privacy-focused media downloaders.Nick: smart.servier.comGuise: runable.comWhy:High-level technical resources; one offers medical illustrations, while the other focuses on executable code environments.Nick:Workout.coolGuise: KouponWhy:Personal optimization tools—one for physical fitness routines and the other for optimizing shopping/savings.Nick: Internet Archive:https://web.archive.org/Guise: Same.newWhy:Part of the ".new" domain movement, providing instant, one-click access to start a new coding or collaborative project.Edtech Throwdown: Vote on twitter @edtechthrowdown and under the pinned post on the profile.Segment 3: Where to Find EdTech ThrowdownDo us a few favors:Subscribe to the Edtech Throwdown PodcastApple PodcastsSpotifyAmazon PodcastsStitcher YouTube Twitter FacebookWrite us an Apple Podcast Review!Tell your friends aboutwww.edtechthrowdown.comTell your friends about the Teach Better Podcast NetworkSubscribe to our Podcast Channels and SocialsApple PodcastsSpotify YouTube Twitter (@edtechthrowdown)FacebookInstagramConnect with us on Social MediaGuise's Social MediaTwitter(@guisegotteched)LinkedInNick's Social...

Chat GPT Podcast
Breaking the AI long context bottleneck

Chat GPT Podcast

Play Episode Listen Later Aug 9, 2026 22:25 Transcription Available


The provided sources describe the development and technical foundations of Llama 2 Long, a series of open-source language models designed to effectively handle extended context windows of up to 32,768 tokens. Researchers from Meta achieved this through continual pretraining on long-form data and a critical modification to Rotary Position Embeddings (RoPE), which reduces the numerical decay that typically hinders a model's ability to process distant information. This approach significantly improves performance on complex tasks like document summarization and long-form question answering while simultaneously boosting results on standard short-context benchmarks. Furthermore, the authors introduce a cost-effective instruction tuning method using synthetic data that allows the model to surpass proprietary alternatives like GPT-3.5-turbo-16k. The documentation also includes a theoretical analysis of positional encoding granularity and validates that these scaling improvements follow a predictable power-law relationship. Consistent with the original Llama 2 series, the models maintain stringent safety standards even when processing much denser information.10 sources

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE 7 de agosto de 2026

Jay Fonseca

Play Episode Listen Later Aug 7, 2026 17:35


PODCAST LAS NOTICIAS CON CALLE 7 de agosto de 2026 - ERP atrasa los pagos del gobierno y no está funcionando - El Nuevo Día Renuncia el presidente de la Junta de la AAA - Jay Fonseca PR Se despide Francisco Domenech en mensaje - Facebook Thommy pide sacar al presidente de la AAA, JGo le dice que no puede cambiar el piloto, TRS le dice que por eso PR está sin agua y luz Congreso vuelve a decirle que no a PR para el programa SNAP - El Nuevo Día Opus Miramar ahora volvería a ser un Condo hotel para poder recibir beneficios contributivos - El Nuevo Día Siempre innovando y con los mejores beneficios, MCS Personal Directo te ofrece cubiertas accesibles para que cuides de tu salud y la de los tuyos.Con una amplia red de proveedores de más de 15,000 médicos de libre selección. Reembolso de hasta $40 mensuales por membresía a un gimnasio o por un entrenador personal debidamente certificado. Asistencia en el hogar para servicios de cerrajería, plomería y electricidad de hasta $350 por evento hasta 4 veces al año.¡Únete HOY a la gran familia de MCS!¡Salud que completa tu vida! Llama al 787.945.1259 y oriéntate.Endoso pagado #MCS #incluyeauspicio Trump rescata el peso argentino y el yen japonés - Bloomberg El Fondo del Seguro del Estado reconoce problemas en inversiones y cartera - El Nuevo Día Arranca racionamiento para Carraízo - El Vocero Trump vuelve a traquetear con la ciudadanía americana - WSJGuaynabo soltó 7 millones en contratos sin subasta - Noticel PPD usará Ai para manejar donativos políticos y detectar aportaciones indebidas - Metro Marco Rubio le mete más presión todavía a Cuba - Axios  LOS DATOS DEL DÍA (cierre 6 agosto) Brent~$81/barril (subiendo) Diésel PR$1.25–$1.32/litro S&P 5007,709.56 (−0.2%) Dow53,870 (−0.9%, −479 pts) Bono 10Y4.67% (+0.06) Euro/USD1.1519 (−0.3%) Gas natural$2.67/MMBtu (−0.8%) Hipoteca 30Y6.69%

Noel Díaz - ESNE
Busca, llama y se te abrirá

Noel Díaz - ESNE

Play Episode Listen Later Aug 5, 2026 22:37


Busca, llama y se te abrirá la puerta que anhela tu corazón; el gozo del encuentro con Jesús te inundará y serán colmadas todas tus expectativas, la paz será tu reposo. Sólo debes encontrar a quién siempre te ha esperado, el Padre amoroso #EnLaHroraDelEncuentro

Contado por el Neuropediatra
Ep.4x171 - Lo que hemos aprendido este curso sobre TDAH, Autismo y desarrollo infantil

Contado por el Neuropediatra

Play Episode Listen Later Aug 4, 2026 54:26


Llegamos al último episodio antes de nuestro descanso de verano. Y para cerrar el curso no queríamos hacer un simple repaso de cifras, proyectos o momentos destacados, sino detenernos en algo mucho más útil: todo lo que hemos aprendido escuchando a las familias.Durante estos meses se han repetido muchas dudas sobre TDAH, autismo, lenguaje, conducta, aprendizaje, pantallas o tratamientos. También hemos visto preocupaciones nuevas, mitos que siguen muy presentes y pequeños avances que, a veces, las familias no valoran lo suficiente.Hoy me siento a charlar con Manuel para hacer balance desde esa perspectiva: qué ha marcado este curso, qué nos ha sorprendido y qué deberían tener en cuenta las familias durante el verano. Además, aprovecharemos para responder de forma rápida a algunas afirmaciones que escuchamos constantemente.Con este episodio nos despedimos hasta septiembre, pero antes queremos dejaros información práctica, tranquilidad y algunas ideas para disfrutar del verano sin perder de vista lo importante.¡Dale al PLAY y nos vemos dentro!Si tienes un hijo con algún problema neurológico o sospechas que puede ser la causa de sus dificultades y quieres que te guiemos por el camino correcto, ve ahora mismo a descargar las guías gratuitas para padres que tengo en la web www.elneuropeditara.es. En menos de 15 minutos podrás tener una idea bastante clara de qué le pasa a tu hijo y los pasos a seguir para ayudarle.Si ya tienes claro que valore a tu hijo o quieres una segunda opinión, Ponte en contacto ahora mismo con nosotros para que analicemos tu caso y nos pongamos manos a la obra. Llama al 682 651 047 o escríbenos al mail recepcionista@elneuropediatra.es

PRN - Garage Pass Podcast
Ryan Blaney Has A Pet Llama, And The Story Behind It Is Incredible

PRN - Garage Pass Podcast

Play Episode Listen Later Aug 4, 2026 3:00 Transcription Available


Storypillar
Summertime Is Joke Time Laugh-o-Rama-Llama! 1: Lions, Shovels, and Holey Socks

Storypillar

Play Episode Listen Later Aug 3, 2026 6:43


Summertime Is Joke Time Laugh-o-Rama-Llama! 1: Lions, Shovels, and Holey SocksWe may be taking a mid-summer break… but that doesn't mean your ears have to! Join us for an extra special episode featuring your favorite jokes, a bucket of popsicles and spaghetti, and our new llama friend, Professor Pickles! Featuring jokes from: Kai (Age 7), Harper (Age 10), Jamie (Age 10), Jayna (Age 7), Layla (Age 7), Grace (Age 12), and Professor Pickles (Age 8 in llama years; 32 in human years)Links for Kids: Silly Summer Jokes for KidsWe'll be back with our regularly scheduled Sneak Attacks, Full Episodes, and Brain Breathers on Monday, 9/7/26. Until then, head to storypillar.com and catch up on your favorite episodes.Make a donation! Support Storypillar!https://ko-fi.com/storypillar Shop at: storypillarstore.threadless.comInfo/Get in Touch: Website: www.storypillar.com Instagram: @storypillar Join our mailing list. Created, Written, and Produced by: Meg Lewis Storypillar Theme Song: Lyrics by Meg Lewis Music by Meg Lewis, Andy Jobe, and Suzanna Bridges Produced by Andy Jobe Episode Cover Art: Mackenzie AllisonSound Effects and Additional Music: -https://freesound.org/ -Llama sounds: https://deadsounds.com/lama-sound#google_vignette, https://animalsounds.online/animals/llama -Joke Time Song: https://freesound.org/people/BlondPanda/sounds/659889/ Know a kid with great advice for Sticky Situations? Check out www.storypillar.com/unsticktricks.© 2026 PowerMouse Press, LLC

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

The Future of Everything presented by Stanford Engineering
Best of: The future of AI and the law

The Future of Everything presented by Stanford Engineering

Play Episode Listen Later Jul 31, 2026 34:26


These days, AI is everywhere, and it's increasingly hard to separate the gains from the slop. With that in mind, we're re-releasing my conversation with Stanford Law professor Daniel Ho on the future of AI and the law. When we look for applications where AI can deliver measurable benefit, the legal profession stands out, both for its potential gains in efficiency and equity, and for how much is at stake if we get it wrong. Dan's research — from using AI to identify racist property covenants buried in county deed records, to mapping obsolete regulations that waste thousands of hours of government time — shows what's possible when the technology is applied with rigor and purpose. If you're curious about how AI can serve both justice and good governance, this one is well worth another listen. Have a question for Russ? Send it our way in writing or via voice memo, and it might be featured on an upcoming episode. Please introduce yourself, let us know where you're listening from, and share your question. You can send questions to thefutureofeverything@stanford.edu. Episode Reference Links: Stanford Profile: Dan Ho Connect With Us: Episode Transcripts >>> The Future of Everything Website Connect with Russ >>> Threads / Bluesky / Mastodon Connect with School of Engineering >>> Twitter/X / Instagram / LinkedIn / Facebook Chapters: (00:00:00) Introduction Russ Altman introduces guest Dan Ho, a professor of law, political science, and computer science at Stanford University. (00:02:19) Path into Legal AI How Ho's background shaped his interest in law, and technology. (00:03:35) What Lawyers Do What makes law a complex domain for AI. (00:05:28) Legal Hallucinations When AI  performs well and when it fails.  (00:07:52) Searching Legal Records in California How AI can help identify outdated, harmful, or legally important material. (00:10:28) Scaling Redaction How a model accelerated a process that overwhelmed county recorder offices. (00:13:04) Legal Reform at Scale How AI has supported legal reform by scanning massive bodies of law. (00:15:02) STARA & The City of San Francisco How AI was used to go through San Francisco's code and clean up reporting. (00:20:53) Outdated Obligations How “regulatory sludge” takes the time & resources of the public service (00:25:02) Open vs. Closed AI The differences and associated risks of the different AI systems. (00:30:58) Legal Chatbots Why legal chatbots are promising but risky. (00:33:42) Conclusion Connect With Us:Episode Transcripts >>> The Future of Everything WebsiteConnect with Russ >>> Threads / Bluesky / MastodonConnect with School of Engineering >>>Twitter/X / Instagram / LinkedIn / Facebook Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Iglesia Cristiana Biblica Raah
EP. 64. Daniel: el que busca y llama

Iglesia Cristiana Biblica Raah

Play Episode Listen Later Jul 31, 2026 7:39


Daniel: el que busca y llama Las mejores oraciones casi siempre nacen de una Biblia abierta. En este episodio contemplamos cómo Daniel escuchó primero la Palabra de Dios y luego convirtió sus promesas en oración, descansando únicamente en la misericordia del Señor. Daniel 9:1–19 Escucha este episodio y recuerda que el Padre escucha a quienes se acercan a Él, no por sus méritos, sino por la gracia que ha dado en Jesucristo. __________________________________________________________________________________________________________ Cada domingo escuchamos la Palabra de Dios proclamada. Pero ¿cómo podemos seguir meditándola durante la semana? Te invitamos a escuchar Recordando el Sermón, el devocional diario de la Iglesia Cristiana Bíblica Ra'ah en Bogotá, Colombia. Breves reflexiones, aplicaciones prácticas y momentos de oración basados en la predicación del Día del Señor. Disponible de lunes a viernes.

Josh Bersin
Open Source Models: An Exciting New Business Model For Enterprise AI

Josh Bersin

Play Episode Listen Later Jul 30, 2026 17:07


New names: Kimi K3, Llama, Nemotron, Mistral, Cohere, Deepseek, Phi-4 – these are just a few of the fast-growing open source models from major AI providers. These systems threaten the business models and financial plans of OpenAI, Anthropic, Google, and X.ai. They perform at levels close to Frontier models and the can run up to five-times cheaper on a variety of hardware platforms. What is the disruptive impact of these open source LLMs and how does this impact your AI investments? As you'll hear in the podcast, Open Source unleashes the opportunity for lower cost AI solutions and more vertical, specialized, application-focused solutions we need. And the business model for these systems moves away from the massive investments of the Frontier providers. The result is more complicated than “open means control.” Model tuning, performance, and optimization could be in your future – as AI moves from a platform to a true layered product set we can use as we need. Lots to learn about here, let us know if you have any questions. Additional Information What's the difference between closed, open source, and open-weight AI? A researcher explains What Is Open-Weights A.I.? Comparison of Open Source Models Chapters (00:00:00) - Open Source and the AI Industry(00:11:46) - The Future of AI Is Fully Integrated(00:15:35) - HR 2030

Cabalá: Lecciones Diarias | mp3 #kab_spa
Rabash. ¿Qué significa que el aceite se llama "buenas acciones" en el trabajo?. 32 (1989) [2026-07-30]

Cabalá: Lecciones Diarias | mp3 #kab_spa

Play Episode Listen Later Jul 30, 2026 59:59


Audio, spa_t_norav_2026-07-30_lesson_rb-1989-32-shemen-nikra_n2_p1. Lesson_part :: Daily_lesson 2

Es la Mañana de Federico
La República de los Tonnntos: Sarah Santaolalla se llama idiota a sí misma

Es la Mañana de Federico

Play Episode Listen Later Jul 29, 2026 10:06


Santiago González comenta el balance de Sánchez y la entrevista a Sarah Santaolalla en la que se llama idiota a sí misma.

Live Long and Master Aging
Living Well on Purpose | Emma Magnolia

Live Long and Master Aging

Play Episode Listen Later Jul 29, 2026 32:41 Transcription Available


What if living a longer, healthier life isn't about finding the next supplement or wellness trend, but about making better choices every day?Holistic health educator Emma Magnolia shares her approach to intentional living, exploring how strength training, hydration, sleep, simple daily routines and consistency can help support healthy aging.Rather than chasing quick fixes, this conversation focuses on practical habits that almost anyone can adopt to build a healthier future.Emma Magnolia is the founder of emmawellness.com ----DISCLOSURE: This podcast is supported by affiliate arrangements with a select number of companies. We have arranged discounts on certain products and receive a small commission on sales. The income helps to cover production costs and ensures that our interviews remain free for all to listen. Visit our LIVE LONG SHOP for more details: PartiQlar supplementsEnhance your wellness journey with PartiQlar supplements. No magic formulas, just pure single ingredients, like NMN, L-Glutathione, Spermidine, Resveratrol, TMG and Quercetin. Get a 15% discount with the code MASTERAGING15 at PartiQlarEnergyBits algae snacksA microscopic form of life that could help us age better. Use code LLAMA for a 20 percent discountPartiQlar supplementsEnhance your wellness journey with pure single ingredients. 15% DISCOUNT - use code: MASTERAGING15SiPhox Health home blood testingMeasure 17 critical blood biomarkers from home. Get a 20% discount with code LLAMA Disclaimer: This post contains affiliate links. If you make a purchase, I may receive a commission at no extra cost to you.Support the showThe Live Long podcast, a HealthSpan Media LLC production, shares ideas but does not offer medical advice.  If you have health concerns of any kind, or you are considering adopting a new diet or exercise regime, you should consult your doctor.

El Larguero
Entrevista | Joan García no llama a la Champions una "obsesión", pero sí una "exigencia": "Este año seguro que vamos a estar más cerca"

El Larguero

Play Episode Listen Later Jul 29, 2026 7:21


Joan García, reciente campeón del mundo con la selección, atiende a El Larguero desde su campus de porteros en Tordera (Barcelona). Mientras el conjunto blaugrana prepara la temporada 2026/27 en el centro de alto rendimiento de la Federación Inglesa, St. George's Park, el guardameta del equipo que ha conseguido la segunda estrella y pieza clave en el de Hansi Flick, valora en El Larguero el post Mundial y las expectativas el equipo culé... con la Champions muy presente. 

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE DE 28 DE JULIO

Jay Fonseca

Play Episode Listen Later Jul 28, 2026 15:59


PODCAST LAS NOTICIAS CON CALLE DE 28 DE JULIO - 1200 millones menos entre fondos federales y presupuesto de PR para el próximo año fiscal - El Vocero Trump dice estar impresionado con Zelesnky y su capacidad militar - WSJHoy se reúne Trump con Netanyahu en la visita del primer ministro - NYTLa Fed decide mañana: ¿y si SUBE las tasas?Un momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequías que se aproximan. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado. Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#incluyeauspicio#windmarhome  Trump va a la Corte Suprema para impedir el voto por correo - CNNRacionamiento inminente en Carraízo y sus clientes - El Vocero Investigan casos de hospital por agua asquerosa - Primera Hora CRIM busca dueños de 55 mil propiedades que no aparecen - El Nuevo Dia Fonalledas apoyan a Jenniffer y dice que le dan la bienvenida a las primarias - El Vocero Alegan que Cosculluela llamó a joven para amenazarla por estar con otros tipos y la amenazó con matar a su familia - El Vocero Viva la ley de plásticos de un solo uso y todavía investigan si la van a implementar o no - El Vocero No sabemos qué hacer con el sargazo en PR - Primera Hora Alcaldes defienden cobro de impuestos a fondos federales - El Nuevo Día Menos protección para animales en peligro de extinción - El Nuevo Día No cuadran los números del fondo de desempleo, aparenta haber montones de fraudes - El Nuevo Día Entidades falsas creando estudiantes fatuos para cobrar becas Pell - El Nuevo Día Juramenta nueva presidenta hoy en Perú. Keiko Fujimori y la derecha conquista Latinoamérica - El Nuevo Día Trump quiere que MAHA le meta mano a eliminar las vacunas para niños - WSJEl SAVE Act no tiene los 60 votos, Trump exige aprobarlo sí o sí - Punchbowl News PR importó $3,254 millones en genéricos en 2025, por lo que los aranceles le darán oportunidad y tumbe a la vez -  LOS DATOS DEL DÍA (cierre lunes 27 jul) Brent≈ $83/barril · cae fuerte por pausa Irán-EEUU Diésel (retail EEUU)a la baja siguiendo al crudo (dato aprox.) S&P 5007,413.18 · +0.02% Dow Jones52,210.08 · +0.51% Nasdaq24,932.08 · -0.18% Bono 10 años≈ 4.65% Euro/USD1.1397 Gas natural$2.72/MMBtu · -1.75% Hipoteca 30 años6.58% (Freddie Mac) / ~6.75% (Bankrate)

Pod and Prejudice
Mansfield Park Volume 3 Chapters 8-9

Pod and Prejudice

Play Episode Listen Later Jul 28, 2026 68:20


Portsmouth is worse than Fanny could've imagined, and in these chapters, she tries to find a friend among her family. Mary sends a gossipy letter to Fanny, and Fanny kind of likes it. Fanny buys a knife for her little sister AND joins a library! Topics discussed include regency era dentistry, Fanny's llama mama, Mrs. Price's resemblance to her sisters, the social status of the Ward sisters, Spicy Susan, the Saturday half-holiday, the comparison of Portsmouth to celibacy, Mary poking the bear, Baron Wildenhaim's courtship of Julia, what qualities in a person are innate, and luxurious and daring wealth.Patron Study Questions come from Avi, Kate, Ghenet, and Linnea. Topics discussed include how Fanny's sisters could've benefitted if she'd come home earlier, Mrs. Price as a tragic figure, and Fanny's lack of a female friendship, where Fanny truly belongs, whether Portsmouth is an enticement or warning about marrying Henry, and Maria as a warning.Becca's Study Questions: Topics discussed include the picture Jane Austen is painting of the lower classes and how Susan complicates that picture, what role Susan plays in the narrative, whether Fanny really misses Mansfield, why Maria is unhappy, and predictions about our former main characters. Funniest Quote(s):“As to the little irritations, sometimes introduced by Aunt Norris, they were short, they were trifling, they were a drop of water to the ocean, compared with the ceaseless tumult of her present abode.”“I hope she will recollect it and be satisfied, as well she may, with moving the queen on a palace, though the king may appear best in the background.”Questions moving forward: How close will Fanny and Susan get? Will we go to the Rushworths' party? Will Edmund propose?Who wins the chapters? Mary Crawford!Glossary of People, Places, and Things: Arthur (Mr. Ratburn and the Special Someone), High School Musical, Is Your Mama a Llama?, Mean Girls, Rugrats, the White Lotus, Dr. Johnson, the Office, Money with MelNext Episode: Mansfield Park Volume III Chapters 10-11 or Chapters 41 and 42Our show art was created by Torrence Browne, and our audio is produced by Graham Cook. For bios and transcripts, check out our website at podandprejudice.com. Pod and Prejudice is transcribed by speechdocs.com. To support the show, check out our Patreon! Check out our merch at https://podandprejudice.dashery.com.Instagram: @podandprejudiceTwitter: @podandprejudiceFacebook: Pod and PrejudiceYoutube: Pod and PrejudiceMerch store: https://podandprejudice.dashery.com/

Asticharlas con Julio Astillero
Lunes 27 de julio de 2026 | ¿Error o burla? Trump llama a México "Me-hee-ho" durante un discurso

Asticharlas con Julio Astillero

Play Episode Listen Later Jul 28, 2026 29:54


¿Error o burla? Trump llama a México "Me-hee-ho" durante un discursoEnlace para apoyar vía Patreon:https://www.patreon.com/julioastilleroEnlace para hacer donaciones vía PayPal:https://www.paypal.me/julioastilleroCuenta para hacer transferencias a cuenta BBVA a nombre de Julio Hernández López: 1539408017CLABE: 012 320 01539408017 2Tienda:https://julioastillerotienda.com/ Hosted on Acast. See acast.com/privacy for more information.

Noticentro
SEP llama a frenar uso excesivo de redes sociales en menores

Noticentro

Play Episode Listen Later Jul 27, 2026 1:48 Transcription Available


Sheinbaum destaca respaldo a limitar celulares en planteles educativos Protección Civil activa alerta por lluvias intensas en el paísIrán endurece postura frente a EE. UU.Más información en nuestro podcast#grc

Noticentro
Línea 6 del Metrobús opera con desvíos

Noticentro

Play Episode Listen Later Jul 25, 2026 1:51 Transcription Available


Hallan restos humanos abandonados en Temixco, Morelos Renuncia ministro de Educación de India por filtración de exámenesDía Naranja llama a erradicar la violencia contra mujeresMás información en nuestro podcast#grc

RNZ: Country Life
Addicted to yarn: Llama herd follows love of fibre

RNZ: Country Life

Play Episode Listen Later Jul 24, 2026 13:43


It was the 1990s when many farms were looking to diversify. 'Why not llamas?', thought Southland spinner and farmer Janette Buckingham.You can find photos and read more about the stories in this episode on our webpage, here.With thanks to:Janette Buckingham, Thickthorne LlamasGo to this episode on rnz.co.nz for more details

RNZ: Country Life
FULL SHOW: Country Life for 24 July 2026

RNZ: Country Life

Play Episode Listen Later Jul 24, 2026 50:58


This week Country life talks to an American poultry farmer about her bird flu experience, visits a llama farmer and meets a family who've been farming celery for more than a century.You can find photos and read more about the stories in this episode on our webpage, here.In this episode:0:59 - Rural news wrap5:45- What New Zealand can learn from a farmer who lived through H5N117:43 - Addicted to yarn: Llama herd follows love of fibre31:36 - Celery, butterflies and Travis the goatWith thanks to:Georgie Cartanza, University of DelawareJanette Buckingham, Thickthorne LlamasGraham Franklin, Lucy Franklin, Alan Franklin, Monique Franklin, Jasmine Franklin, Luke Franklin, Brian King, and Solola Manulele, Franklin FarmMake sure you're following us on your favourite podcast app, so you don't miss new episodes every Friday evening.Send us your feedback or get in touch at country@rnz.co.nzGo to this episode on rnz.co.nz for more details

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan

Growthaholics
#321 - Por que a Apple paga US$ 1 bi por ano pro Google (e a Meta não consegue copiar) | Growthaholics

Growthaholics

Play Episode Listen Later Jul 23, 2026 47:23


Google, Apple e Meta: a guerra da inteligência artificial que vai decidir o futuro dos negócios e do mercado de trabalho. Neste episódio de Por Trás do Hype, Pedro Waengertner recebe André Lopes, repórter de tecnologia da Exame, para contar os bastidores da reportagem que mapeou toda a estratégia de IA do Google desde 2017. Você vai entender por que a Apple paga bilhões para usar o Gemini na Siri, por que a Meta não consegue destravar o Llama, o que o chip próprio da OpenAI significa na corrida contra a Nvidia, e por que só 17% do mundo realmente usa inteligência artificial no dia a dia. Um raio-x direto ao ponto sobre tecnologia, startups e o futuro do trabalho, pra quem empreende, lidera times ou quer sair na frente.Nos siga nas redes sociais: Pedro WaengertnerACE VenturesEXAME

Contado por el Neuropediatra
Ep.4x170 - Al otro lado del teléfono: todo lo que ocurre antes de entrar en consulta

Contado por el Neuropediatra

Play Episode Listen Later Jul 22, 2026 27:50


Cuando una familia contacta por primera vez con el centro, muchas veces no llega únicamente buscando una cita. Llega con dudas, con preocupación, con prisas o sin saber muy bien cuál debe ser el siguiente paso.Antes de entrar en consulta hay una parte fundamental del proceso que no siempre se ve: escuchar, orientar, organizar agendas, coordinar gestiones y procurar que cada persona se sienta atendida desde el primer momento.Hoy me siento a charlar con May Jiménez, que forma parte del equipo de recepción y está en contacto directo con nuestros pacientes y sus familias. Queremos descubrir cómo es realmente ese trabajo, qué ocurre detrás de cada llamada y cuántas piezas hay que mover para que todo parezca sencillo cuando llegáis al centro.¡Dale al PLAY y nos vemos dentro!Si tienes un hijo con algún problema neurológico o sospechas que puede ser la causa de sus dificultades y quieres que te guiemos por el camino correcto, ve ahora mismo a descargar las guías gratuitas para padres que tengo en la web www.elneuropeditara.es. En menos de 15 minutos podrás tener una idea bastante clara de qué le pasa a tu hijo y los pasos a seguir para ayudarle.Si ya tienes claro que valore a tu hijo o quieres una segunda opinión, Ponte en contacto ahora mismo con nosotros para que analicemos tu caso y nos pongamos manos a la obra. Llama al 682 651 047 o escríbenos al mail recepcionista@elneuropediatra.es

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE DE 21 DE JULIO

Jay Fonseca

Play Episode Listen Later Jul 21, 2026 19:44


PODCAST LAS NOTICIAS CON CALLE DE 21 DE JULIO -  Regresan las multas de AutoExpreso - El Vocero Hutíes plantean cerrar paso a barcos de Arabia Saudita disparando el precio del petróleo a 90 - BBC Miguel Romero pide sacar a Itza García y a Francisco Domenech - TeleOnce Nuevo operador para inspeccionnes de carros otra vez empieza el proceso - El Vocero Brasil le pasa a USA como exportador de alimentos mundial - Semafor  Quieren limitar la ciudadanía americana a nacidos en territorios de USA - El Nuevo Día Acuden al tribunal federal para ayudar a que PR no tenga que pagar la deuda y advierten que es falso que afectemos al mercado municipal de bonos - El Nuevo Día Un momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequías que se aproximan. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado. Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#windmarhome#windmar #incluyeauspicio Se disparan los impagos en préstamos estudiantiles, PR lidera todo USA  - El Nuevo Dia Frutos exóticos en San Sebastián, boricua siempre montones de frutas que se dan en PR que nadie imaginaría que se dan aquí - Primera Hora Carraízo, Cidra, Matrullas y Loco entran a embalses de observación por falta de lluvia - Primera Hora Hogar de envejecientes ceierra tras estar operando sin permiso - WAPA Federales arrestan sujeto por estar amenazando a pentecostales, judíos, soldados americanos y otros - Noticel Muere Granados Navedo, exvicepresidente de la Cámara bajo el PNP - Metro Trump vuelve a meterle aranceles a Canadá de 50%, papas y varios productos LOS DATOS DEL DÍABrent~$89/barril · tocó $90 WTI~$83/barril Gasolina EE.UU. (retail)>$4.00/galón (+13¢ semana) S&P 5007,443.28 (−0.19%) Dow Jones51,839.26 (−0.59%) Nasdaq25,508.07 (−0.05%) Bono 10 años4.59% Gas natural (Henry Hub)~$3.70/MMBtu (prom. 2026, EIA)Euro/USD e hipoteca 30 añossin confirmar hoyCierre del 20 de julio. En PR (DACO, mediados de julio): gasolina regular ~$1.05-$1.10/litro, diésel ~$1.23-$1.32/litro.

Resilient Cyber
Resilient Cyber w/ Joshua Saxe - Why Restricting AI Makes Us Less Secure

Resilient Cyber

Play Episode Listen Later Jul 20, 2026 37:12 Transcription Available


Does restricting frontier AI in the name of safety actually make us less secure? Joshua Saxe joins me to make the case that it does, and that AI cybersecurity will be won through defender adoption, not restriction.Josh has spent 15 years at the intersection of AI and security. He built and ran the machine learning program at Sophos, then led security for Llama at Meta, covering security post training, evals, agent guardrails, and prompt injection prevention. He recently left to co-found a startup reimagining vulnerability and exposure management agentically. He also writes one of the most cited blogs on AI and cyber policy.In this episode:- Why restricting frontier model access harms defenders more than attackers- How monitored closed models put threat actors at a structural disadvantage- The jagged frontier, and why attackers don't need frontier models for most of their tradecraft- The national security and supply chain risks of pushing the world onto Chinese open weights models- Why exploits don't cause cyberattacks, and which attacker constituencies AI actually unblocks- The dual use ceiling on guardrails and classifiers- Where defenders should be adopting AI right now, from access management to SOC automation- Using agents to burn down the mountain of security technical debtChapters:0:00 Intro0:42 Josh's background, from blackhat teen to Llama security lead3:07 The case for diffusion over restriction6:14 Why restriction hurts defenders more than attackers10:19 The jagged frontier and what attackers actually use models for12:49 National security and the supply chain risk of Chinese open weights16:08 Exploits don't cause cyberattacks20:20 Where defenders should adopt AI right now24:20 Guardrails, classifiers, and the dual use problem27:34 Reimagining vulnerability management with agents32:17 The structural advantage defenders hold35:15 Policy wishes and the attacker's Claude Code momentFollow Josh:LinkedIn: https://www.linkedin.com/in/joshua-saxe-01845a1Substack: https://joshuasaxe181906.substack.comFollow Resilient Cyber:Substack: https://www.resilientcyber.ioSubscribe for more conversations with security practitioners and leaders.#aisecurity #cybersecurity #vulnerabilitymanagement #aipolicy #opensourceai

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE DE 17 DE JULIO

Jay Fonseca

Play Episode Listen Later Jul 17, 2026 15:23


PODCAST LAS NOTICIAS CON CALLE DE 17 DE JULIO - Mega proyecto para arreglar en varias fases el super acueducto - El Nuevo Día500 mil de fianza a sujeto que iba a dejar morir o matar viejita en Las Piedras - El Vocero Se supone que al fin acabe el escaneo de furgones - El Nuevo Día Gas Natural de USA el gran ganador de guerra de Ormuz - Semáforo Trump jura que CHINA interviene en elecciones de USA y comete fraude - CNN Corrección podría mudarse por renta muy cara - El Nuevo Día Supremo no decide si va a resolver el caso de LUMA de arranque - El Nuevo Día Guerra de grandes ligas por pago de COBRA tras impuestos a la construcción de alcaldes - El Nuevo Día Siempre innovando y con los mejores beneficios, MCS Personal Directo te ofrece cubiertas accesibles para que cuides de tu salud y la de los tuyos.Con una amplia red de proveedores de más de 15,000 médicos de libre selección. Reembolso de hasta $40 mensuales por membresía a un gimnasio o por un entrenador personal debidamente certificado. Asistencia en el hogar para servicios de cerrajería, plomería y electricidad de hasta $350 por evento hasta 4 veces al año.¡Únete HOY a la gran familia de MCS!¡Salud que completa tu vida! Llama al 787.945.1259 y oriéntate.Endoso pagado#incluyeauspicio #MCSEmpezarán a cobrar el tren urbano otra vez - El. Vocero Educación dice que ha cumplido tanto con los padres de educación especial que merece que bajen multa de 11 mil diarios a mil - El Nuevo Día Demócratas inundados en dinero en comparación con republicanos para midterms - Punchbowl News Fin de semana sin IVU para escuelas arranca hoy - El Nuevo Día Buscan apoyo a proyecto de status - El Nuevo D´â Ahora Valerie Rodz dice que Domenech sí intervino para detener solicitud de propuestas - El Nuevo Día Le tendrán que dar a los fondos buitres de la deuda de la AEE expediente para ver si gastaron dinero que era para la deuda - El Nuevo DíaLOS DATOS DEL DÍA  Brent ~$84.93/barril · cerca de máximos de 1 mesDiésel/gasolina al alza siguiendo el crudo (dato PR sin confirmar)S&P 500 7,533.77 (−0.51%)Dow Jones 52,552.97 (−0.20%, −105.67 pts)Bono 10 años 4.59% (+0.04)Euro/USD 1.1433 (−0.27%)Gas natural Henry Hub sin confirmar; gas europeo €55/MWh (máx. desde marzo)Hipoteca 30 años 6.55%

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE DE 16 DE JULIO

Jay Fonseca

Play Episode Listen Later Jul 16, 2026 21:12


PODCAST LAS NOTICIAS CON CALLE DE 16 DE JULIO - Gobierno dice que LUMA tiene que irse en el Supremo de PR - El Vocero 200 millones para rehabilitar el tren urbano - El Vocero Nuevo arancel de 25% a Brasil entra el 22 de julio; escalada comercial en curso.Los demócratas se rompen por Israel, sobre 100 votan contra ayudar a Israel - Semafor Bonistas logran aliado en Rivera Schatz, alega que la Junta solo quiere ofrecer poco para seguir quedándose en PR y no negociar pa guisar en PR - Noticel AAA dice que dará créditos por falta de servicio - El Vocero Trump tiene mensaje especial esta noche a las 9PM Un momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequías que se aproximan. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado. Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#incluyeauspicio #windmarhomeAlcaldes se enteran por los medios de racionamiento de agua, Comerío, Cidra, Canóvanas y Río Grande, San Lorenzo no tendrá, no anunciaron en Carolina - El Vocero EEUU atacó un petrolero y objetivos en el norte de Irán; Teherán respondió contra bases estadounidenses en Baréin, Kuwait y Jordania, y advierte que resistirá "hasta el final".Líderes de inteligencia artificial piden regulación de la inteligencia artificial - Axios Al fin aprueban ley para prohibir el escaneo obligatorio de furgones - El Vocero LePen está al frente en las encuestas de Francia - Reuters Se harán 3 casinos más en PR - El Vocero No hay jurisdicción federal en caso de Domenech y Sebastián hasta ahora - El Nuevo Día Huevos bajan demasiado de precio tras dumpeo federal - El Nuevo Día Albergue se va a quedar con perro que atacó a secretario de Recursos Naturales - El Nuevo Día Corrección no paga la renta de su sede y quieren contrato de 4 millones por quedarse en oficinas - El Nuevo Día Tres testigos dicen que Elvia sacó algo punzante de cartera - El Nuevo Día Sigue la guerra de Paulson v. Ghaffar en el tribunal de PR - El Nuevo Día Este weekend es el back to school sin IVU   LOS DATOS DEL DÍA (cierre / referencia 15-16 jul) Brent$84.95 · +12% en 3 sesiones  Gasolina EEUU (AAA)$3.89/gal · +9¢ semana  S&P 500 (fut.)~7,610 · -0.1%  Dow (fut.)+145 pts · +0.3%  Bono 10 años4.57% Euro/USD1.1431 · +0.1% Gas natural (Europa)€53.1/MWh  Hipoteca 30 años~6.65%

Live Long and Master Aging
Hidden Drivers Of Longevity | Oscar Trelles

Live Long and Master Aging

Play Episode Listen Later Jul 13, 2026 37:08 Transcription Available


Modern life asks less and less of our bodies, while placing ever greater demands on our minds. We move less, sleep poorly and fill every spare moment with stimulation—often without realizing the long-term consequences.Oscar Trelles explores the connections between recovery, resilience and the way we age. His work—through the wellness and performance company Breathing Flame—focuses on helping people better understand the conditions that shape health and long-term wellbeing.In his forthcoming book, The Human OS Manual, he argues that a longer, healthier life depends less on isolated interventions and more on the rhythms and routines that shape our days.So have we lost touch with the conditions that help us thrive—and what would it take to restore them?----DISCLOSURE: This podcast is supported by affiliate arrangements with a select number of companies. We have arranged discounts on certain products and receive a small commission on sales. The income helps to cover production costs and ensures that our interviews remain free for all to listen. Visit our LIVE LONG SHOP for more details: PartiQlar supplementsEnhance your wellness journey with PartiQlar supplements. No magic formulas, just pure single ingredients, like NMN, L-Glutathione, Spermidine, Resveratrol, TMG and Quercetin. Get a 15% discount with the code MASTERAGING15 at PartiQlarEnergyBits algae snacksA microscopic form of life that could help us age better. Use code LLAMA for a 20 percent discountDisclaimer: This post contains affiliate links. If you make a purchase, I may receive a commission at no extra cost to you.Support the showThe Live Long podcast, a HealthSpan Media LLC production, shares ideas but does not offer medical advice.  If you have health concerns of any kind, or you are considering adopting a new diet or exercise regime, you should consult your doctor.

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE DE 10 DE JULIO

Jay Fonseca

Play Episode Listen Later Jul 10, 2026 19:51


PODCAST LAS NOTICIAS CON CALLE DE 10 DE JULIO - Israel dice que Irán iba a asesinar a Trump y por eso cambiaron de avión presidencial - CNN 22 personas y empresas se declaran en quiebra por día - El Nuevo Día Gobernadora entrega 200 títulos de propiedad - NEWSPR Acusan a Gian Carlo Piovanetti por vivir como rico cogiendo de tonto a clientes y fraude - El Nuevo Día Plantean darle alivios a ayunadores y cuidadores, costo de 300 millones - El Nuevo Día  PR es el líder en enfermedades raras en todo USA - El Vocero Ciencias Forenses busca encontrar personas desaparecidas para dar paz a familias que no encuentran a sus seres queridos - El Vocero Politank dice que se va a defender de todo esto sal pa fuera - El Vocero Agricultura y comida en aumento de precio por sequía y fertilizantes - El Nuevo Día Nombran fiscal investigadora en caso de Negrón Reichard - El Nuevo DíaUn momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequías que se aproximan. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado. Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#windmarhome #incluyeauspicioPetróleo/diésel: Brent ~$76; diésel de EE.UU. al alza más rápida en 4 años; Rusia prohíbe exportar diésel (≈30% de su refinación estuvo fuera el mes pasado)Segundo día de ataques Irán–EE.UU.OpenAI y Google vendieron modelos avanzados a subsidiarias en Singapur de Alibaba, Baidu y Tencent — empresas en la lista negra del Pentágono - FT  LOS DATOS DEL DÍA (snapshot Bloomberg, 10 jul) Brent$76.13 (-0.2%) S&P 500 (futuros)7,580.75 (-0.1%) Nasdaq 100 (futuros)29,817.25 (-0.4%) Bono 10 años4.53% (-0.02) Oro$4,104.95 (-0.5%) Diésel EE.UU.alza más rápida en 4 años (nivel s/c)

Big Technology Podcast
Meta CTO Andrew Bosworth: Our Path To Frontier AI, Renting Models, Consumer AI's Struggles

Big Technology Podcast

Play Episode Listen Later Jul 8, 2026 46:52


Andrew "Boz" Bosworth is the chief technology officer of Meta. Bosworth joins Big Technology to discuss why Meta fell behind in the frontier AI race and how it plans to turn its models, products, and distribution into an advantage. Tune in to hear his candid explanation of what went wrong with Llama, why the best AI products will use multiple models, and what it will take for consumer agents to break through. We also cover Meta's AI glasses, the future of augmented reality, employee tracking and training programs, AI companions, and the painful process of adapting a company to a technological revolution. Hit play for a revealing conversation about Meta's AI comeback and the products that could shape how we interact with computers. --- Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice. Watch the full documentary here: https://www.gravitee.io/ai-agent-documentary Want a discount for Big Technology on Substack + Discord? Here's 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b Learn more about your ad choices. Visit megaphone.fm/adchoices

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE DE 3 DE JULIO

Jay Fonseca

Play Episode Listen Later Jul 3, 2026 22:39


- 118 escuelas bilingües en PR Gobierno dice tiene superávit de 635 millones No hay break para sacar a la Junta dicen los expertos - El Nuevo Día Legislatura propone que CRIM tenga que publicar propiedades por embargar para que se puedan poner en el mercado de compra de vivienda - El Vocero Informe federal advierte que reconstrucción eléctrica de PR está demasiado segmentada - Noticel Vivienda contrata a jefe de OGPE bajo FEI - Jay Fonseca PR Siempre innovando y con los mejores beneficios, MCS Personal Directo te ofrece cubiertas accesibles para que cuides de tu salud y la de los tuyos.Con una amplia red de proveedores de más de 15,000 médicos de libre selección. Reembolso de hasta $40 mensuales por membresía a un gimnasio o por un entrenador personal debidamente certificado. Asistencia en el hogar para servicios de cerrajería, plomería y electricidad de hasta $350 por evento hasta 4 veces al año.¡Únete HOY a la gran familia de MCS!¡Salud que completa tu vida! Llama al 787.945.1259 y oriéntate.Endoso pagado#mcs#incluyeauspicio Abel Nazario niega gestiones con agricultor y cuadrarle reunión con gobernadora - Noticel Menos conserjes escolares, supuestamente 40% menos - El Vocero La semana que viene se va jefe de fiscalía federal - El Nuevo Día Sujeto trató de meter 356 pastillas de Suboxone a la cárcel federal - El Nuevo Día Más cargos contra sujeto que era asesor legislativo y a la vez empresario por los donativos legislativos a quebrada Margarita - El Nuevo Día Guerra por contratos de seguridad en el gobierno - El Nuevo Día Irán despide hoy al Ayatollah Ataques de Rusia contra Ucrania dejan 30 muertos - Reuters Gobernadora pide explicaciones la Junta de libertad bajo palabra - El Vocero Archivada querella contra Ferraiuoli - El Vocero Canadá tendrá sus NBA en juego que PR no tendrá a Alvarado - Metro  2,295 muertos por terremotos de Venezuela Se espera mega bajón de precios del petróleo, Citi dice que a 60 el barril - Bloomberg Peligroso químico en aeropuerto de Aguadilla Se casa Taylor Swift y Travis Kelce - Washington Post  • ⁃ Trump chotió plan de Israel para matar nuevo liderado de Irán 

Jay Fonseca
PODCAST LAS NOTICIAS CON CALLE DE 2 DE JULIO

Jay Fonseca

Play Episode Listen Later Jul 2, 2026 18:51


PODCAST LAS NOTICIAS CON CALLE DE 2 DE JULIO - Rusia importa gasolina y combustibles de India, mientras ataca bestialmente a la capital de Ucrania - Financial Times Otro caso de libertad bajo palabra de asesino suelto, la policía no tenía ni registro ni archivos de los casos que había tenido el tres veces asesino  - El Vocero 280 policías menos en PR, 176 por retiro, 13 por retiro obligatorio, 62 renuncias Gobernadora dice que tenemos superavit de 635 millones y que salen estados financieros auditados del 2023 - El VoceroSacan de abanderado oficialmente a Tuto Bermúdez - Metro Un momento para WindMar Home — la empresa con más de 20 años protegiendo los hogares puertorriqueños.Solar para bajar tu factura. Techo para proteger tu inversión. Agua para que nunca te quedes sin — especialmente con las sequías que se aproximan. Y batería para total independencia energética.Todo bajo una misma empresa. Un solo llamado.Llama al 787-489-1155 o visita windmarhome.comWindMar Home — los que se preparan hoy , duermen tranquilos mañana.#windmarhome#incluyeauspicioArranca rehabilitación con 170 millones para Sergio Cuevas, mientras dicen que eso durará 4 años - El Vocero OpenAi propone darle 5% al gobierno federal como dueños de ChatGPT - FTCRIM fabricó documentos para justificar el gasto de jangueo que nunca se dio y evento que nunca ocurrió - El Vocero   Junta le impuso presupuesto de la UPR y son 566 millones, 421 millones menos que en 2017 - El Nuevo Día Sueltan 91 millones de FEMA, pero ponen nuevo requisito para soltarnos dinero federal - Metro Hermana de Gabriela Nicole vio que Elvia entregó objeto a AnthonieskaVuelven a quedar en nada negociaciones con Irán - Reuters El precio de la carne se dispara al tener menos ganado en reserva en 75 años - Reuters AEE demanda a dueños de Genera porque tuvieron que gastar 54 millones más por no traernos suficiente Gas Natural y tener que usar diesel - El Nuevo DíaTrump se sale del acuerdo con Canada y México que é mismo había establecido - Economist Sony anuncia el fin del disco físico y de ahora en adelante todo será digital - Bloomberg  Llegan a Venezuela primeros rescatistas y suministros desde PR - Wapa Federales van a realizar arrestos en Jardines de Loíza y Loiza Home for the Elderly, pero FBI niega operativo - WAPA Trump hizo 2.2 billones de billetes - Bloomberg GAO dice que a PR no le han soltado casi nada para arreglar sistema energético - El Nuevo Día  LOS DATOS DEL DÍA (cierre 1 jul 2026) Brent~$71-72/bbl ▼ mín. desde feb. · WTI ~$67.57 Diésel wholesale~$3.21/gal (aprox.) S&P 5007,483.23 −0.22% Dow Jones52,305 −0.03% Bono 10 años~4.49% Euro/USD1.1407 Gas natural (Henry Hub)~$3.19/MMBtu (aprox.) Hipoteca 30 años~6.5%

Chente Ydrach
GIOVA LE TIRA A ARTE CARDE, PALESTINO SE RETIRA Y KELE LLAMA A GALLO

Chente Ydrach

Play Episode Listen Later Jun 19, 2026 73:31