Podcasts about KV

  • 1,408PODCASTS
  • 5,215EPISODES
  • 41mAVG DURATION
  • 2DAILY NEW EPISODES
  • Aug 28, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about KV

Show all podcasts related to kv

Latest podcast episodes about KV

Pratt on Texas
Episode 4053: Trump approval now low in Hispanic S Texas districts he won big | Wildfire update – Pratt on Texas 8/28/2026

Pratt on Texas

Play Episode Listen Later Aug 28, 2026 43:33


The news of Texas covered today includes:Our Lone Star story of the day: U.S. House Speaker Mike Johnson was in Texas yesterday campaigning for Republican Congressional nominees and said that the Trump Agenda is on the ballot. But that may indeed be the problem with the redistricted seats where Trump won big two years ago, today Trump's approval rating is terrible in these areas. Link to poll.Our Lone Star story of the day is sponsored by Allied Compliance Services providing the best service in DOT, business and personal drug and alcohol testing since 1995.Lubbock ISD trustees vote to raise property taxes on citizens instead of cutting spending.Wildfire update. Ross Fire scorches 85,000+ acres, second-largest in North Texas history (we haven't been able to accurately track the size of these things for very long.)Oil and gas rig count rises this week in Texas but only by one.The famed Texas Rangers, oldest profession law enforcement force in North America, get a new chief.Federal judge dismisses Fort Worth Muslim's lawsuit over principal reassignment.PUC approves routes for first 765-kV power lines in Texas despite much objection.Listen on the radio, or station stream, at 5pm Central. Click for our radio and streaming affiliates.www.PrattonTexas.com

BardsFM
Mike Adams: China Wins the AI Race & the Case for Sovereign Cognition │ BardsFM

BardsFM

Play Episode Listen Later Aug 25, 2026 59:22


Episode 4210 │ August 25, 2026 Anthropic admitted degrading their own AI's output quality to comply with EU watermarking law. Mike Adams says that changed the entire race. WHAT THIS EPISODE COVERS  Scott Kesterson sits down with Mike Adams for their first conversation in a year, opening with the oracle versus librarian framework both have independently arrived at — AI as a tool for research and task execution, never as a divine verdict-giver — and using the Chinese farmer Wu's destroyed sesame crop as the case study for why blaming AI rather than human accountability misdiagnoses every failure. Adams delivers the most significant claim of the interview: Anthropic publicly admitted to watermarking all model outputs under EU law by statistically biasing token selection, a technical concession that provably reduces output quality, while China's DeepSeek, Qwen, and other open-published models have pulled ahead of the US frontier labs specifically because Chinese companies shared their research openly instead of hoarding it as classified defense technology. The conversation closes on the sovereign computing model both are independently building — desktop AI on AMD's Strix Halo chip, open-weight models run locally on solar power, and a coming economy where owning hardware and generating your own energy matters more than money, because decentralized cognition is, in Adams's words, the single greatest threat to globalist control of what people are allowed to know. KEY QUESTIONS ADDRESSED  What did Anthropic publicly admit about watermarking their AI model outputs to comply with EU law — and why does Mike Adams argue this technical concession provably reduces the quality of every response, even for US domestic users who never consented to it? How did China overtake the US in the AI race — and why does Adams argue that China's strategy of openly publishing research papers and sharing innovations like sparse attention mechanisms and KV cache compression across companies proved more effective than the secrecy-driven US model treating AI as classified defense technology? What is the sovereign computing model Scott and Mike Adams are both building — and why does running open-weight AI models locally on solar-powered hardware represent, in Adams's words, decentralized augmented cognition and the single greatest threat to globalist controllers who have always sought to limit public access to knowledge? ABOUT BARDSFM BardsFM is a daily independent podcast covering faith, liberty, history, and information warfare. Hosted by Scott Kesterson — combat veteran, documentary filmmaker, and rancher. Over 4,100 episodes and 50 million lifetime downloads. New episodes every weekday. bards.fm This episode was researched and produced under the Spatial Terra Intelligence Methodology (STIM v5) — the analytical framework built by Scott Kesterson — with AI-assisted research synthesis at a 70/30 human/AI authorship ratio, fully disclosed. All analysis, conclusions, and editorial judgments are those of Scott Kesterson. BardsFM's archive includes hundreds of episodes on prayer, scripture, and walking the Way of Christ — available free in the full episode catalog. DOWNLOADS Citizen's Guide - Community Organizing Against Data Centers: click here Citizen's Guide - Auditing Automatic License Plate Readers: click here Citizen's Guide - Auditing Your State's Driver License Data: click here AFFILIATE LINKS Bards Nation Health Store: www.bardsnationhealth.com MYPillow promo code: BARDS >> Go to https://www.mypillow.com/bards and use the promo code BARDS or... Call 1-800-975-2939.  EMPShield protect your vehicles and home. Promo code BARDS: Click here Treadlite Broadforks...best garden tool EVER. Promo code BARDS26: TreadliteBroadforks.com EnviroKlenz Air Purification, promo code BARDS to save 10%: www.enviroklenz.com Morning Intro Music Provided by Brian Kahanek: www.briankahanek.com Founders Bible 20% discount code: BARDS >>> TheFoundersBible.com Windblown Media 20% Discount with promo code BARDS: windblownmedia.com White Oak Pastures Grassfed Meats, Get $20 off any order $150 or more. Promo Code BARDS: www.whiteoakpastures.com/BARDS Mission Darkness Faraday Bags and RF Shielding. Promo code BARDS: Click here DONATIONS: If you wish to support this podcast directly you can donate here... DONATE: Click here MAILING ADDRESS: Xpedition Cafe, LLC Attn. Scott Kesterson 591 E Central Ave, #740 Sutherlin, OR  97479

Karlovy Vary
Náš host: Festival na ulici se koná v DEPU2015. Náměstí chybí, ale snažíme se nebrečet, říká programový šéf

Karlovy Vary

Play Episode Listen Later Aug 17, 2026 16:03


Kvůli rekonstrukci náměstí Republiky v Plzni se 33. ročník Festivalu na ulici přesunul do DEPA2015 a jeho blízkého okolí.

Náš host
Festival na ulici se koná v DEPU2015. Náměstí chybí, ale snažíme se nebrečet, říká programový šéf

Náš host

Play Episode Listen Later Aug 17, 2026 16:03


Kvůli rekonstrukci náměstí Republiky v Plzni se 33. ročník Festivalu na ulici přesunul do DEPA2015 a jeho blízkého okolí.Všechny díly podcastu Náš host můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Medierna
Mycket på spel när Radio Aftonbladet drar igång

Medierna

Play Episode Listen Later Aug 15, 2026 29:51


Och så frågar vi oss hur mycket droger Sommar i P1 tål och fördjupar oss i den tuppfäktning som uppstår när en journalist intervjuar en annan journalist. Lyssna på alla avsnitt i Sveriges Radios app. Kvällstidningen bryter ny markPå måndag morgon är det dags. Radio Aftonbladet ger sig ut på outforskad mark, när den röda sändningslampan tänds på Sveriges första reklamfinansierade radionyhetskanal. Många höjde på ögonbrynen redan när det stod klart att Aftonbladet i slutet av förra året gav sig in i kampen om ett av de tre så åtråvärda, och numera väldigt mycket billigare, nationella sändningstillstånden. Är analog radio verkligen rätt framtidshäst för mediehuset att satsa på? Mediemyndigheten verkade i alla fall tro på Aftonbladets vision och gav dem NRJ:s tidigare tillstånd.Så nu är det alltså upp till bevis.Vår reporter Edvin Ek åkte ut till den nystartade redaktionen och togs emot av Cajsa Lindberg, konsult, programledaren Alex Letic och tidningens publisher Lotta Folcker.Joel Kinnaman fick inte lägga ut texten om ”plant medicine”Efter en lugn och skandalfri säsong fick Sommar i P1 äntligen igång en liten snackis nu i veckan, när den svenske Hollywoodskådisen i programmet berättade att Sveriges Radio stoppat honom från att dela med sig av sina positiva drogupplevelser. Rimligt eller fjantigt?Joanna Korbutiak ringde upp programmets ansvarige utgivare Emma Boëthius och Leonidas Aretakis, chefredaktör för Flamman och författare till boken Extas i folkhemmet – Sveriges psykedeliska historia.Robinsonkritiken som fick Treutiger att tappa tålamodetI den femte och sista delen av vår sommarserie Dålig stämning i studion tar vi med dig tillbaka till sommaren 1997, och ett klassiskt exempel på den tuppfäktning som lätt kan uppstå när en journalist ställs mot en annan journalist.Robin Jonsson gör sitt bästa för att mäta sig med programledaren och journalisten Harald Treutiger.

Dobré dopoledne
Květiny na svatbě nemusí stát statisíce. Floristka radí, kam peníze opravdu dát

Dobré dopoledne

Play Episode Listen Later Aug 14, 2026 25:36


Květiny dokážou svatbě dodat nezapomenutelnou atmosféru. Floristka Markéta Jašová z Rostlinopis Třebíč ale upozorňuje, že některé dekorace mohou být zbytečným finančním zatížením.Všechny díly podcastu Dobré dopoledne můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Plus
Věda Plus: Mikrovlnná trouba a opravy silnic

Plus

Play Episode Listen Later Aug 13, 2026 26:39


Jak na první pohled nelogické spojení úspěšně funguje na našich cestách. - O několik týdnů napřed jsou stromy v českých lesích. Kvůli suchu a horku začaly už teď shazovat listí. Lesníci sčítají i ztráty na jarních sazenicích, Kolik procent mladých stromů uhynulo třeba v brněnských městských lesích? - Máte rádi kávu? Vědci už mají po desetiletích sporů jasno - je nebo není pro zdraví škodlivá? - Co všechno by se mohlo vyrábět z mořských řas? A můžou nahradit materiály vyráběné z fosilních zdrojů? Moderuje Štěpánka Martanová.

BIT-BUY-BIT's podcast
What a Week | THE BITCOIN BRIEF 86

BIT-BUY-BIT's podcast

Play Episode Listen Later Aug 12, 2026 61:48 Transcription Available


A bi-weekly news show informing you on the latest in Bitcoin, privacy and open source tech hosted by Ungovernables, Max and Q. THIS IS THE TWEET Q WANTS YOU TO SEE: https://x.com/justh0dl/status/2086202393998291138AOBWhat a fucking weekKeyOS v1.3.1 now publicly availableSomething exciting to share on Friday's FTF (delayed by 2 weeks)NEWSThe Coldcard entropy catastropheSources: Coinkite technical backgrounder / The Rage, L0la L33tz / TRM Labs / coldcard.rip / cktripwire.com / Bitcoin Magazine victim surveyEXPLAINERThe largest self-custody theft on record, and it traces back to a single wrong conditional. In March 2021 a build guard checked whether Coldcard's hardware random number generator was defined rather than whether it was enabled, so seed generation silently fell back to a deterministic software PRNG. Every seed made on an affected device from that point carried roughly 40 bits of entropy on Mk2 and Mk3, and about 72 on Mk4, Mk5 and Q, instead of the intended 128. That is guessable. Someone did the maths offline, derived the addresses, checked them against the public chain, and swept everything with a balance. Somewhere between 1,400 and 1,800 bitcoin gone, depending on whose forensics you trust, with a median victim loss of one BTC. The part people keep missing: updating the firmware does not fix an existing seed. A weak seed is weak forever.ACTION FOR LISTENERSMove funds to a brand new seed BEFORE upgrading firmware (Lopp's guidance, on reports of update problems).Use high fees. If you see your own coins in the mempool, the attacker opted into RBF and you can outbid them. Window is minutes.Multisig users: consider a private mempool like Marathon's Slipstream.Keep the device; the UID may prove ownership in any recovery process.Updating does NOT fix an existing seed. A weak seed is weak forever.You are exempt only if you added 50+ fair independent private dice rolls, or used a strong unique BIP39 passphrase stored separately.BTCPay Server: unauthenticated LND macaroon theft, actively exploitedSources: BTCPay security advisory / v2.4.2 release / CoinDesk / TFTCCRITICAL FRAMING NOTE: this is ONE story, not two. The "BTCPay bug" and the "LND credential exploit" are the same event. The vulnerability is the macaroon leak. The Aug 8-9 wave of coverage is follow-up hardening, not a new incident. Do not present them separately.REMEDIATION (updating alone is NOT enough)Update to v2.4.2. Verify "2.4.2" in the footer.Update NBXplorer to 2.6.10+.Revoke and regenerate LND macaroons. Updating stops new theft but does nothing about already-stolen credentials. Deleting files is insufficient; the macaroon root signing key must be destroyed at node level. v2.4.2 does this automatically for standard Docker deployments. Custom reverse proxies, separate Tor services or port forwarding must rotate manually.Move funds out of any BTCPay-generated on-chain hot wallet and recreate it.Update LND to 0.21.1. Audit for unrecognised channel closures, unknown peers, unexplained balance changes.SIDE EFFECT WORTH FLAGGING: v2.4.2 removes public LND API access on Docker deployments, which breaks remote wallet connections such as Zeus connecting to your own BTCPay node. Intentional, no restoration timeline published.BREAKING CHANGE: Greenfield Basic authentication disabled by default five minutes after account creation (#7492). BTCPay: "We are not aware of any user impacted by this breaking change, as API Keys authentication is generally used."Boltz suspends all swaps indefinitelySources: Boltz statement / canary.boltz.exchange / The Defiant / TFTC / Bull Bitcoin statementCORRECTION TO THE COMMON FRAMING: Boltz has not shut down. It suspended swap services indefinitely. And the canary sequence runs the opposite way to the rumour: lapsed → suspended → renewed clean.The Bitcoin Red TeamSources: Calle and Rob Hamilton on Nostr/X / Bitcoin Magazine / CoinDesk / TFTC / OpenSats Red Team FundSUMMARYThis is the story that explains the other four. After the Coldcard exploit, Calle and Rob Hamilton pointed frontier AI models at the open-source Bitcoin stack and started auditing everything. In 108 hours, 25 developers scanned 501 projects and produced 7,958 findings, 1,280 of them rated high or critical, at a compute cost north of 58,000 dollars. They found the BTCPay bug's neighbours, and Boltz cited exactly this dynamic when it switched itself off. The uncomfortable symmetry is that the same capability doing the defending is what an attacker almost certainly used on Coldcard in the first place. And the bottleneck turns out not to be finding bugs, it is telling anyone: only 19.5% of the projects they scanned even have a SECURITY.md file, and only 13.1% list a security contact. The scanners move at machine speed. Responsible disclosure is still hunting around for an email address.BIP-110: the fork that mined two blocks and frozeSources: bip110monitor.com / Peter Todd code review / Aaron van Wirdum, Bitcoin Magazine / Lopp's Layman's Guide / Saylor essay / CoinDeskRELEASESBitcoin core / protocollibsecp256k1 v0.8.0 - 2026-08-03Adds a native Silent Payments (BIP-352) module directly into the crypto library nearly every self-custody wallet builds on, plus up to ~11% faster signature verification. Quietly the most consequential positive release of the fortnight: it lowers the bar for every wallet to ship reusable static receive addresses.Bitcoin Knots v29.4 - 2026-08-08Non-urgent maintenance: fixes a chainstate DB bug causing repeated large rewrites, and adds corruption-detection safeguards around BIP-110 mandatory signaling. No critical fixes. (No Bitcoin Core release in window; latest is v31.1 from 2026-07-08.)Hardware / signingColdcard Firmware 4.2.0 (Mk2/Mk3) - 2026-08-03The patch for the entropy catastrophe. Affected ranges: Mk2/Mk3 4.0.1 through 4.1.9; Mk4/Mk5 all before 5.6.0; Q all before 1.5.0Q. Companion fixes shipped the same day: 5.6.0 Mk4/Mk5, 1.5.0Q, 6.6.0X Edge, 6.6.0QX Edge Q. Updating does NOT fix an existing seed - changelog says Mk3 users "must regenerate any seeds made on earlier versions as their entropy is critically low at just ~40 bits." TAPSIGNER, OPENDIME and SATSCARD unaffected.Krux 26.08.0 - 2026-08-04Maintainer odudex is stepping down and the project may be archived. "Krux was not created by me: Jeff started it and passed it on to me, and now it is my turn to pass the torch." On succession: "Krux may be carried on by another maintainer, if a proof-of-work backed Krux contributor accepts the role. Otherwise the Krux project will be put in sunset mode and gracefully archived in a few months." Cause is hardware, not drama: "K210 chips are no longer produced, and Canaan dropped the Kendryte line entirely." Substantial release regardless: fixes a heap buffer overflow in the camera entropy module, adds stricter PSBT fee-calculation checks, replaces the Python UR stack with a faster C module, and makes Krux Installer fully offline.Frostsnap v0.3.0 - 2026-08-05FROST threshold-signing device ships reproducible/deterministic builds and "a fresh release signing key as part of an overhauled release-signing pipeline." Well timed in a fortnight where "can you verify what is running on your signer" is the whole conversation. Catch: the new key breaks in-place Android updates, so direct-APK users must uninstall, reinstall, and re-visit their threshold devices to restore.Trezor Suite v26.7.4 - 2026-08-04Lowers minimum Normal-priority fee rate to 0.2 sat/vB and ships updated Safe 7/5/3 and Model T firmware with security improvements.BitBoxApp 4.51.4 - 2026-08-07Bundles new BitBox02 firmware v9.26.5.Specter Desktop v2.1.11 - 2026-08-09Genuinely security-relevant: adds auth and CSRF protection to the HWI bridge settings, restores validation of active API tokens so revoked JWTs are rejected, and warns that Specter's auth layer does not encrypt the data folder. Also ships an in-app Coldcard Mk3 seed-entropy advisory.Bitkey source/2026-08-02-0031 - 2026-08-02Block's consumer hardware wallet, routine source drop.LightningBTCPay Server v2.4.2 - 2026-08-07Actively exploited, funds already stolen. "This release contains fix of a critical vulnerability that is being actively exploited. You need to update as fast as you can." Unauthenticated remote .macaroon disclosure for LND, plus a TOTP 2FA bypass via Greenfield Basic auth. Requires NBXplorer 2.6.10. Breaking change: Basic auth disabled by default five minutes after account creation. See News item 2 for full remediation.lnd v0.21.2-beta.rc1 and v0.20.3-beta.rc1 - 2026-08-08Not security releases and not related to the BTCPay exploit. Migration/stability fixes only: KV-to-SQL payment migration edge case, channeldb migration recovery, invoice handling, data races, bounded memory on graph sync.Zeus v13.1.3 - 2026-07-27Adds LND v0.21.1-beta support for embedded and remote nodes; patches known vulnerabilities in the ws, js-yaml and markdown-it dependencies. (Note: Zeus also shipped an unreleased swap-security sprint on 08-04 - verify…

VOV - Sự kiện và Bàn luận
Tiêu điểm - Tuyên Quang: Quyết tâm hoàn thành trường nội trú liên cấp đúng tiến độ

VOV - Sự kiện và Bàn luận

Play Episode Listen Later Aug 12, 2026 4:34


VOV1 - Chỉ còn ít ngày nữa sẽ bước vào năm học mới. Với thầy và trò ở vùng sâu, vùng xa, biên giới của tỉnh Tuyên Quang, việc sớm hoàn thành các trường phổ thông nội trú liên cấp tiểu học và trung học cơ sở không chỉ là mong mỏi, mà còn là điều kiện để học sinh có một môi trường học tập tốt hơn.Điểm trường phổ thông nội trú liên cấp TH&THCS Xín Mần nằm ở độ cao hơn 1.700 mét so với mực nước biển. Đây có thể được xem là một trong những điểm trường ở độ cao nhất cả nước. Những ngày này, mưa liên tục tại khu vực xã Xín Mần và dọc tuyến biên giới của tỉnh Tuyên Quang. Trên quốc lộ 177, 178, nhiều đoạn đường vào xã Xín Mần xuất hiện sạt lở. Đường vốn đã xuống cấp, nay lại càng khó đi do mưa kéo dài.Thời tiết không thuận lợi khiến nhiều hạng mục của công trình phải tạm dừng. Trong khi đó, ngày khai giảng đã cận kề, áp lực về tiến độ ngày càng lớn. Anh Lê Việt Hùng, cán bộ kỹ thuật Công ty TNHH một thành viên Đông Bắc, chia sẻ:Bây giờ công trình thi công vẫn còn gần một nửa, hiện tại gặp rất nhiều khó khăn. Thời tiết mưa gió, nhân lực thi công hạn chế, vận chuyển vật liệu vậtt tư nhiều chỗ đá lở, đường lở nên vật liệu lên đến đây rất khó khăn.Không chỉ khó khăn trong vận chuyển vật liệu, nguồn nước phục vụ thi công cũng là một bài toán nan giải. Nghịch lý là khi trời mưa, nước nhiều thì không thể thi công, lúc trời nắng, điều kiện thi công tốt lại không có nước, đơn vị phải tìm cách đưa nước từ nơi khác về. Anh Lã Chí Trường, Công ty TNHH một thành viên Đông Bắc, cho biết: Cũng tính toán làm đường nước vừa để sinh hoạt vừa để thi công, phải chạy một đường ống khoảng 9 km lấy nước về. Nước kiểm tra thì cơ bản đảm bảo tạm đáp ứng, thi thoảng đường ống bị lỗi phải đi sửa chữa.Điểm trường phổ thông nội trú liên cấp TH&THCS Xín Mần gồm 10 hạng mục, trong đó có 8 khối nhà chính, một hạng mục hạ tầng kỹ thuật và một hạng mục phụ trợ. Đến thời điểm này, việc triển khai dự án đã chậm so với kế hoạch đặt ra. Nguyên nhân chủ yếu là trong quá trình san nền xuất hiện sạt lở, phải xử lý hiện trường, mất nhiều thời gian. Cùng với đó, mạch nước ngầm phát sinh tại khu vực thi công làm gia tăng nguy cơ sạt lún nền móng, mái taluy, đồng thời ảnh hưởng đến 2 hộ dân và 2 vị trí cột điện đường dây 35 Kv. Địa chất phức tạp cũng khiến phương án thi công phải liên tục điều chỉnh. Ông Lã Chí Quân, Phó Giám đốc Công ty TNHH một thành viên Đông Bắc, cho biết:Với địa hình địa chất trên này rất phức tạp, tất cả các khối nhà đều phải thay đổi kết cấu móng lý do là đều nằm trên vùng địa chất đá phong hóa, cứ đào ra là sạt, hiện tại vẫn đang bị sạt. Chúng tôi đang cùng chủ đầu tư, đơn vị tư vấn thiết kế nghiên cứu phương án để xử lý.

Liberec
Zprávy pro Liberecký kraj: Pumptrack pro cyklisty i koloběžkáře. Semily investují do prostoru pro náctileté, přistaví i skatep…

Liberec

Play Episode Listen Later Aug 8, 2026 2:01


Kvůli vysoké poptávce po prostoru pro trávení volného času náctiletých bude mít projekt i další etapy. Přímo v bazénu má vzniknout nový skatepark a obepínat ho bude in-line dráha.

kv kraj kolob libereck pumptrack semily
Svenska Mordhistorier
Försvinnandet i Smedby

Svenska Mordhistorier

Play Episode Listen Later Aug 7, 2026 30:24


I det här avsnittet vill vi varna för att det förekommer beskrivningar av grovt våld mot barn. Kvällen den 29 mars 1999 kliver en man i fyrtioårsåldern in på polishuset i Norrköping. Han vill anmäla sin sambo och deras gemensamma spädbarn, försvunna. Den orolige mannen tror att de kan ha åkt utomlands, på en kryssning till Finland. Men de har inte kommit tillbaka och inte heller hört av sig på flera dagar. I själva verket vet mannen mer än han låtsats om och återvända kommer familjemedlemmarna aldrig att göra. Det här är ett gammalt avsnitt från Podme. För att få tillgång till Podmes alla premiumpoddar samt fler avsnitt från den här podden, helt utan reklam, prova Podme Premium kostnadsfritt.

men finland kv norrk podme podme premium podmes
Texas Ag Today
Texas Ag Today - August 6, 2026

Texas Ag Today

Play Episode Listen Later Aug 6, 2026 23:53


We're still waiting to rebuild the cow herd. The Farm bill failed on a first vote in Senate Ag Committee.There are alternatives to 765 kV transmission lines.The cattle herd in Hale County is on the rise.A new study shows calves like being with their buddies.

Kvällspasset i P4
Kvällspasset med Adrian Alebo Andersson: Mitt världsarv

Kvällspasset i P4

Play Episode Listen Later Aug 6, 2026 29:44


UNESCO:s världsarvslista är på mångas läppar och nu samlar Kvällspasset ihop lyssnarnas världsarv. Det listas smultronställen, personer och saker! Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Micke nominerar ABBA till Kvällspassets världsarvslista som tack för musiken, Bror slår hårt för att surströmmingen ska ta en plats och Gunnar tycker att ”världens vackraste häst”, dalahästen, ska rätt in på listan! Dessutom hör vi Martin som vill lyfta Sveriges kanske mest okända världsarv – Struves meridianbåge.

Proactive - Interviews for investors
First Phosphate secures additional $4.84M in Canadian Government support

Proactive - Interviews for investors

Play Episode Listen Later Aug 5, 2026 4:10


First Phosphate Corp. CEO John Passalacqua joined Steve Darling from Proactive to discuss the company's latest federal funding milestone after securing $4.84 million in non-repayable contributions from the Government of Canada through Natural Resources Canada's First and Last Mile Fund to advance the Bégin-Lamarche Phosphate Project in Quebec. Passalacqua said the new funding builds on the $16.7 million previously awarded by NRCan in March 2026 through the Global Partnerships Initiative, further strengthening government support for developing the company's high-purity phosphate deposit. Approximately $3.07 million will fund studies related to power transmission infrastructure, including site selection, feasibility work, environmental assessments, public and Indigenous consultations, and the design of a 161-kV transmission line and substations in the Saguenay–Lac-Saint-Jean region. A further $1.77 million will support planning for transportation infrastructure, including preparatory work for a new mine access road and studies evaluating upgrades to regional bypass roads connecting the Bégin-Lamarche project to rail infrastructure and the Port of Saguenay. The work will include technical, environmental, and economic studies, along with traffic analysis and community consultation. Management believes the additional funding will accelerate key infrastructure planning while supporting the long-term development of the Bégin-Lamarche project as a strategic source of high-purity phosphate for the North American lithium iron phosphate (LFP) battery supply chain. #proactiveinvestors #firstphosphatecorp #cse #phos #otcqx #frspf #frspf #BeginLamarche #LFPBatteries #CriticalMinerals #Phosphate #QuebecMining #CleanEnergy #EnergyTransition #MiningNews #canada #governmentofcaanda #NaturalResourcesCanada'sFirstandLastMileFund

Forbes Česko
K-pop jako továrna na sny i zlomené lidi. Sedlák a Mišík chystají sérii K-Dream

Forbes Česko

Play Episode Listen Later Aug 4, 2026 24:29


Režisér Adam Sedlák společně s Adamem Mišíkem a producentkami Monikou Soukup a Lindou Krejčí připravují sérii K-Dream. Pojednává o Evropanovi, který je posedlý představou, že se stane hvězdou k-popu. Kvůli svému snu je ochotný vzdát se čehokoli, i sama sebe.O sérii K-Dream mluví kreativní tým v dalším díle speciální podcastové série Forbes Life KVIFF Talents, která letos vznikla přímo v Císařských lázních v Karlových Varech. Zapojené projekty se tam představují filmovým profesionálům, producentům, partnerům i investorům.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

Ptám se já
Není možné, aby jeden stát ohrožoval ostatní, říká europoslankyně k migrantům v Ceutě

Ptám se já

Play Episode Listen Later Aug 3, 2026 34:08


Po nedávné migrační krizi ve španělské Ceutě čelí Madrid kritice dalších zemí EU, že nezvládá migraci. V neděli se vyjádřila i marocká vláda, která z celé situace obvinila pašeráky a dezinformace. Jak by měla nyní unie postupovat? Hostem Ptám se já byla europoslankyně Nikola Bartůšek (zvolena za Přísahu). O migrační krizi kolem španělského města Ceuta, které leží na africkém pobřeží a do něhož se ve čtvrtek nelegálně dostalo zhruba 60 tisíc migrantů z Maroka, budou v úterý jednat ministři vnitra EU. K mimořádné schůzce vyzvali prezidenti a premiéři dvaadvaceti států unie. Mimo jiné německý kancléř Friedrich Merz, italská premiérka Giorgia Meloniová, polský premiér Donald Tusk i český premiér Andrej Babiš (ANO). K situaci se v neděli poprvé vyjádřila i marocká vláda. Podle tamního ministerstva vnitra byly nedávné události v Ceutě důsledkem kombinace faktorů: akcí sítí pašeráků a převaděčů, dezinformací na sociálních sítích i nesprávné interpretace verdiktů soudu. To vytvořilo u tisíců Maročanů přesvědčení, že se na evropské území mohou dostat snadno a bez právních důsledků. Pašeráckým mafiím už v pátek přičetl migrační krizi v Ceutě také španělský premiér Pedro Sánchez. Zločinecké sítě podle něj využily verdikt španělského nejvyššího soudu, který začátkem července dospěl k závěru, že pohraničníci nemohou okamžitě vyhostit ty, kteří na území města připlavou. Obratem vrátit mohou podle soudu jen toho, kdo překonal hraniční překážku. Kvůli tomu také španělští pohraničníci v sobotu instalovali u Ceuty v moři bariéru z bójí. V pátek už také španělské město opustila většina z několika desítek tisíc běženců, na místě jich ale stále zůstává ještě několik tisíc. Při pokusu dostat se do Ceuty zemřely desítky lidí. Kdo za náhlou migrační vlnou v Ceutě stál? Selhalo Španělsko? A jak zranitelná je Evropská unie náhlými uprchlickými vlnami? -- Podcast Ptám se já. Rozhovory s lidmi, kteří mají vliv, odpovědnost, informace. Sledujte na Seznam Zprávách, poslouchejte na Podcasty.cz a ve všech podcastových aplikacích. Archiv všech dílů najdete tady. Své postřehy, připomínky nebo tipy nám pište prostřednictvím sociálních sítí pod hashtagem #ptamseja nebo na e-mail: audio@sz.cz.

Magiska Godnattsagor
Tomten som blev allergisk mot snö

Magiska Godnattsagor

Play Episode Listen Later Aug 2, 2026 20:54


I dagens avsnitt reser vi till januari 2023 – och landar på en plats som känns konstigt bekant! Kvällens saga "Tomten som blev allergisk mot snö" är önskad av Alve, 7 år från Borås – poddens allra första sagoönskan någonsin.Tomten Totte börjar nysa så fort det snöar, så han flyttar till en solig liten by vid havet i Spanien. Men kan man verkligen sluta sakna sitt gamla hem? En varm saga om att man kan ha plats för två hem i hjärtat – nu i en finare, varmare nyversion.I studion får gänget se sina yngre, nervösa jag spela in poddens allra första avsnitt, och tidsrese-Aida möter sin egen första version med den gamla skrapiga robotrösten. Plus fem fakta om AI och robotar – visste du att ordet "robot" är över hundra år gammalt?Stötta podden och få tillgång till nya sagor! Gå med i Magiska Godnattsagor-klubben!Skicka in förslag på kommande sagor via www.magiskagodnattsagor.seSökord: magiska godnattsagor, godnattsaga, barn, läggdags, podcast för barn, barnlitteratur, ai, godnatt, robot, tomten, jul, poddens början, Borås

Kvällspasset i P4
Rasmus på radioluffen: Abisko

Kvällspasset i P4

Play Episode Listen Later Jul 31, 2026 37:19


Radioluffens sista dag är här! Ikväll hittar vi Rasmus i Abisko nationalpark och vi hör lyssnarnas berättelser om vad de har varit med om där. Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Varje dag besöker Rasmus och Kvällspasset en ny stad! Radioluffen tar med lyssnarna på en resa med start i Karlstad, vidare till Gävle, förbi Umeå, upp till Luleå och avslutas i Abisko.Nu har radioluffen tagit sig till sista stoppet. I Abisko guidas Rasmus av Rickard runt i nationalparken. Vi hör också Sofia som avslutade fjällvandringen med en dusch i ”Frippes fall”, Inger hade sitt drömbröllop på fjället och Lotta med familj kämpade mot en magsjuka under hela vandringen.

Kvällspasset i P4
Rasmus på radioluffen: Luleå

Kvällspasset i P4

Play Episode Listen Later Jul 30, 2026 37:45


Vi har nått radioluffens näst sista stopp och hemma hos Ulrika pratar vi om samiska berättelser, anekdoter och associationer! Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Varje dag besöker Rasmus och Kvällspasset en ny stad! Radioluffen tar med lyssnarna på en resa med start i Karlstad, vidare till Gävle, förbi Umeå, upp till Luleå och avslutas i Abisko.I Luleå har Ulrika med husdjur tagit emot Rasmus med öppna armar. Hos dem får Rasmus bland annat testa kaffedrycken kåsa (eller Guksi som det heter på samiska). Vi pratar också med Per-Erik som berättar om en magisk naturupplevelse, Tilda vars konst har inspirerats av sina samiska rötter och Elmo som åkte till Jokkmokk för att lära sig om samisk kultur och renskötsel. Dessutom ringer Ulrikas son Samuel, mer känd som Robinson-Nutti, och berättar hur han har det på fjället.

Kvällspasset i P4
Rasmus på radioluffen: Umeå

Kvällspasset i P4

Play Episode Listen Later Jul 29, 2026 29:21


Radioluffen har nått Umeå och vi efterlyser lyssnarnas mest oväntade, roligaste eller finaste historier därifrån! Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Varje dag besöker Rasmus och Kvällspasset en ny stad! Radioluffen tar med lyssnarna på en resa med start i Karlstad, vidare till Gävle, förbi Umeå, upp till Luleå och avslutas i Abisko.I Umeå hänger Rasmus hos Magnus där det käkas palt och spelas gitarr. Vi pratar även med Mikael som ger fisketips, Per råkade göra inbrott i en butik och Magnus fru Maggan ringer in för att se vad som händer där hemma när hon är borta.På grund av tekniska problem är ljudkvalitén något nedsatt de första 6 minuterna av podden.

Kvällspasset i P4
Rasmus på radioluffen: Gävle

Kvällspasset i P4

Play Episode Listen Later Jul 28, 2026 37:22


Radioluffen tuffar vidare och idag har Rasmus landat i Gävle! Därför efterlyser vi lyssnarnas mest oväntade historier därifrån. Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Varje dag besöker Rasmus och Kvällspasset en ny stad! Radioluffen tar med lyssnarna på en resa med start i Karlstad, vidare till Gävle, förbi Umeå, upp till Luleå och avslutas i Abisko.Rasmus har kommit hem till Linnea och Ulf i Gävle. Hos dem vankas grill och tiokamp! Vi pratar också med Urban som klippt Börje Salming – utan att veta det, Carina fick en nygammal vän när hon flyttade dit och när Elisabeth var tio år cyklade hon tio mil till Gävle för att besöka sin bror.

Tech Disruptors
Solidigm on AI Driving New Era for Data Storage

Tech Disruptors

Play Episode Listen Later Jul 27, 2026 28:56


Co-CEO of Solidigm Xin Guo joins Bloomberg Intelligence's Jake Silverman on this episode of the Tech Disruptors podcast to discuss how AI is reshaping data-center storage as NAND flash-memory prices continue their unprecedented rise. They explore the technologies driving greater solid-state-drive adoption, including KV cache, which stores AI models' working memory during inference, and why the shift to agentic AI could lift NAND content to as much as 25 exabytes (25 million terabytes) per gigawatt of deployed compute. The conversation also examines how longer-term supply contracts may reshape memory cyclicality, creating a different industry cycle over time.

Kvällspasset i P4
Rasmus på radioluffen: Karlstad

Kvällspasset i P4

Play Episode Listen Later Jul 27, 2026 31:59


Rasmus åker ut på radioluff genom Sverige! Resan börjar i Karlstad och vi hör lyssnarnas berättelser apropå staden. Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Varje dag besöker Rasmus och Kvällspasset en ny stad! Radioluffen tar med lyssnarna på en resa med start i Karlstad, vidare till Gävle, förbi Umeå, upp till Luleå och avslutas i Abisko.I Karlstad fick Inger sitt första sommarjobb, Sören höll i spelningarna med Sven-Ingvars på Sandgrundsudden och Nettan minns Karlstad som ungdomens epicentrum. Dessutom pratar vi med Robinsonvinnaren Sonja Rudqvist som agerar turistguide!

True Story
Kvæleren fra bakkerne 5:5

True Story

Play Episode Listen Later Jul 27, 2026 32:02


Los Angeles er nu en by i knæ, og Kvæleren fra bakkerne er blevet et navn, alle kender og frygter. Mens Angelo Buono forsøger at fastholde kontrollen, begynder Kenneth Bianchi at bevæge sig væk fra Californien og ind i en ny tilværelse, hvor hans fejltrin får afgørende betydning. I Washington sker det gennembrud, som Frank Salerno og efterforskerne har ventet på, og langsomt samles trådene mellem de brutale drab i bakkerne og mændene bag dem. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Osobnost Plus
Chyběly nám stany a domluva s venezuelskou armádou byla složitá, říká k záchranné misi hasič Hädler

Osobnost Plus

Play Episode Listen Later Jul 24, 2026 26:23


S následky ničivého zemětřesení ve Venezuele pomáhal na začátku července speciální tým hasičů z Česka. Žádné živé se sice českým hasičům najít nepodařilo, ale vyprostili pět mrtvých, které mohli jejich příbuzní důstojně pohřbít. „Kvůli vlhku a vedru jsme se museli střídat i po patnácti minutách, velitel nás musel hlídat a stahovat,“ popisuje práci v sutinách a v ochranných kombinézách člen USAR týmu a host Osobnosti Plus Raphael Hädler.Všechny díly podcastu Osobnost Plus můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Názory a argumenty
Jiří Leschtina: Vláda nezařadila SPD do zprávy o extremismu? Marná snaha o únik před realitou

Názory a argumenty

Play Episode Listen Later Jul 24, 2026 4:09


V minulých dnech vláda schválila Zprávu o extremismu a předsudečné nenávisti za rok 2025, do níž ministerstvo vnitra nezahrnulo vládní SPD. Navzdory tomu, že právě loni státní zástupce obžaloval toto hnutí i Tomia Okamuru za šíření nenávisti. Kvůli xenofobním předvolebním plakátům, na nichž byl i muž tmavé pleti, třímající zakrvácený nůž.Všechny díly podcastu Názory a argumenty můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Learning Bayesian Statistics
The Next Step Beyond LLMs: Foundation Models for Inference

Learning Bayesian Statistics

Play Episode Listen Later Jul 22, 2026 5:35


Today's clip is from episode 161, featuring Luigi Acerbi. In this conversation, Luigi explains one of the biggest engineering bottlenecks facing transformer-based probabilistic models—and how his group found a way around it.The core challenge is that many inference models treat data as an unordered set, making them naturally permutation invariant. That's statistically elegant, but computationally painful: every time a new data point arrives, the model has to recompute attention over the entire dataset from scratch, preventing the kind of KV caching that makes modern language models so efficient.Luigi walks through his team's solution: a hybrid architecture that keeps the original context fully set-based while introducing a causal-attention buffer for newly arriving data. The result is dramatically faster inference- up to 100× faster in some settings - opening the door to applications like reinforcement learning, active data acquisition, and, ultimately, Luigi's long-term vision of a foundation model for Bayesian inference.Get the full discussion hereSupport & Resources→ Support the show on Patreon→ Bayesian Modeling Course (first 2 lessons free)Our theme music is « Good Bayesian », by Baba Brinkman (feat MC Lars and Mega Ran). Check out his awesome work

Somna med Henrik
Västkaja Östkaja

Somna med Henrik

Play Episode Listen Later Jul 22, 2026 59:55


Hej Somna. Välkommen till den här timman av vad det nu är. Jag står inför ett val, Somna, men när du hör det här har valet redan gjorts. Det är både spännande och skrämmande. Idag ska jag berätta en historia som börjar mitt i handlingen: det är en smäll. Fönster skallrar, fåglar lyfter, och över Södermalm skingras fyra kajor åt var sitt väderstreck, precis som de lovat varandra.Öst-, Väst-, Nord- och Sydkajan, en kvartett som sedan ungdomen har ett löfte: vid första smäll som hörs flyger de åt var sitt håll och samlas sedan igen för att jämföra vad människorna haft för sig. Den här gången var smällen bara en bil som baktände på Folkungagatan, men samtalet dröjer sig kvar länge på taket. Kajorna pratar om att folk har slutat titta upp, att människor har tappat sin nyfikenhet och går med sina gamnackar och tittar ner i telefonen istället för på himlen, på skorna, på varandra.Mitt i alltihop kliver sotaren Gävlert Garnmärnan upp på taket med en flaska prosecco. Han känner igen dem, de där kajorna som alltid sitter och dömer människorna, och till slut öppnar de sina näbbar och pratar med honom. De presenterar sig kort för varandra. Det blir tyst ett tag, lite pinsamt, innan de hittar in på portkoder. Gävlert kan nämligen alla portkoder på Södermalm, och det visar sig att kajorna kan dem också. De skålar i prosecco och säger "nu lever vi". Stämningen är god ända tills Gävlert blir lite för berusad och gör ett övertramp. Kvällen får en vändning ingen såg komma, en vändning som slutar med Gävert som puttas ner av kajorna och landar på en droskhäst. Tiden är en konstig tunna. En sekund föds man, sen är man nunna. Godnatt Somna. Mer från Somna med Henrik: https://somnamedhenrik.se/Mer om Henrik: https://www.henrikstahl.se/Lyssna utan reklam, få extraavsnitt, spellistor med mera på: https://somnamedhenrik.supercast.com/ Hosted on Acast. See acast.com/privacy for more information.

Kvällspasset i P4
Kvällspasset med Sarit Monastyrski: Ut och cykla

Kvällspasset i P4

Play Episode Listen Later Jul 20, 2026 29:32


Kvällspasset är ute och cyklar! Vi efterlyser lyssnarnas mest oväntade, roliga eller härliga berättelser från cykeln. Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Tobias BMX-tur tog tvärstopp när hunden Roj inte ville fortsätta, under en cykeltur fick Susanne fågelbajs i munnen och Hugo cyklade från Landskrona till Malmö för en konsert. Dessutom snackar vi med ”Taximami” Cathrine som berättar hur jobblivet är nu när alla andra är på semester!

Plus
Hovory: Krajanka: Kolumbie se za vlády Petra rozdělila na dva tábory. Velké podvody ve volbách nebyly

Plus

Play Episode Listen Later Jul 17, 2026 23:10


Eliška Krausová odjela v létě 1968 na několik měsíců do Kolumbie zdokonalit se ve španělštině. Kvůli invazi vojsk Varšavského paktu tam ale neplánovaně zůstala a zemi neopustila ani po pádu Sovětského svazu. Jak nyní vypadá tamní politická situace? „Polovina Kolumbijců volila jednoho kandidáta a skoro stejná polovina toho druhého,“ shrnuje dění v pátečních Hovorech Českého rozhlasu Plus pedagožka na univerzitě v Bogotě.

Karlavagnen
Då gjorde jag bort mig rejält!

Karlavagnen

Play Episode Listen Later Jul 17, 2026 55:37


Det är ju så pinsamt att göra bort sig! Kvällen bjuder på dråpliga händelser och vi undersöker varför vi egentligen tycker det är så jobbigt. Lyssna på alla avsnitt i Sveriges Radios app. Vissa gånger vill man bara försvinna, andra, landar man i ett gapskratt och ofta blir det en bra historia i efterhand. Karlavagnen handlar ikväll om att göra bort sig och känna att man bara vill sjunka under jorden när det går fel. Som när man snubblar på en trottoarkant, pruttar i fel miljö eller spiller på sin bordsgranne under en fin middag. Kvällens programledare Sofia Rågenklint kommer själv bjuda på några pinsamma ögonblick och har du också en historia när du gjorde bort dig rejält så ring oss på 020-22 10 30, mejla på karlavagnen@sverigesradio.se eller skriv till oss på Facebook och Instagram. Slussen öppnar kl 21:00 och programmet börjar 21.40

True Story
Kvæleren fra bakkerne 3:5

True Story

Play Episode Listen Later Jul 13, 2026 39:38


En ung kvinde stiger ind i en blå og hvid Cadillac i den tro, at hun har mødt hjælpsomme fremmede, men turen gennem natten fører hende direkte til det gule hus i Glendale. Mens hendes skæbne folder sig ud bag lukkede døre, vokser frygten i Los Angeles, og navnet Kvæleren fra bakkerne begynder for alvor at sætte sig i byens bevidsthed. Samtidig dykker fortællingen ned i Angelo Buonos fortid, hans kvindehad og de mønstre af vold og kontrol, der længe før drabene har præget hans liv. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Názory a argumenty
Viktor Daněk: Odmítáme se vzdát dividendy z míru

Názory a argumenty

Play Episode Listen Later Jul 13, 2026 5:09


Summit Severoatlantické aliance v turecké Ankaře patřil kvůli kompetenčnímu sporu prezidenta a premiéra mezi asi ty vůbec nejsledovanější v Česku. Kvůli vnitropolitickým tahanicím by nám ale neměl uniknout jeho nejdůležitější závěr. Tlak Spojených států na evropské spojence nikterak nepolevil, právě naopak.Všechny díly podcastu Názory a argumenty můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

The Effortless Podcast
Vertical Memory. Shared Memory. Local Memory - The Effortless Podcast – Episode 23

The Effortless Podcast

Play Episode Listen Later Jul 11, 2026 72:49


In this episode of The Effortless Podcast, Dheeraj Pandey sits down with co-host Amit to dissect the dramatic acceleration of AI over the last few months and map out its next major frontier memory.  Moving past prescriptive frameworks and simple prompt-engineering, they unpack how autonomous agents are shifting the industry's focus from "token maxing" to "impact maxing," forcing a complete rethink of computing architecture. The conversation explores how memory within AI agents cannot remain a flat, horizontal file.  Instead, true enterprise intelligence requires a tiered hierarchy of memory spanning episodic, semantic, and procedural layers that mirrors human psychology and classical hardware caching. Drawing a striking parallel between token anxiety and electric vehicle range anxiety, they make the case for a hybrid CPU-GPU future where structured data, governance, and safety rollbacks are critical to preventing autonomous systems from breaking the bank or deleting databases. Key Topics & Timestamps 00:00 – Summer updates and AI's recent "quantum jump". 01:00 – Token maxing vs. impact maxing & autonomous React loops. 03:00 – Model reliability & using Grep, Sed, and Awk for dynamic context. 07:00 – Terminal text-matching tools explained simply. 08:00 – xAI, data center builds, and Neocloud disruption. 10:00 – Cursor's acquisition & the shift to autonomous harnesses. 12:00 – Desktop hurdles: Sandboxing, Docker, and local firewalls. 14:00 – Coding for the "paranoid path" and failure modes. 18:00 – The Core Thesis: Memory as AI's next major frontier. 21:00 – Caching tiers: KV cache vs. CPU/GPU caches and DRAM. 25:00 – Personal vs. enterprise memory: Turning data into goal-oriented meaning. 32:00 – Enterprise memory grammar: Ontology, identity, and work. 41:00 – Psychology of memory: Episodic, semantic, and procedural structures. 45:00 – Hybrid CPU-GPU needs & the EV range anxiety metaphor. 53:00 – Agent safety: Rollbacks, versioning, and transaction protection. 58:00 – Team intelligence: Bringing AI context to Slack and Teams. 1:01:00 – State vs. skill versioning: The derivative of human intelligence. 1:03:00 – Summary: Memory as data reduction & reinforcement learning. 1:09:00 – Final thoughts: Managing atoms vs. bits & the future of labor. Hosts: Amit Prakash – CEO and Founder at AmpUp, former engineer at Google AdSense and Microsoft Bing, with extensive expertise in distributed systems and machine learning. Dheeraj Pandey – Co-founder and CEO at DevRev, former Co-founder & CEO of Nutanix. A tech visionary with a deep interest in AI, systems, and the future of work. Follow the Hosts: Amit Prakash  LinkedIn – https://www.linkedin.com/in/amit-prakash-50719a2/  Twitter/X – https://x.com/amitp42 Dheeraj Pandey  LinkedIn – https://www.linkedin.com/in/dpandey/  Twitter/X – https://x.com/dheeraj  Share Your Thoughts Have questions, comments, or ideas for future episodes?

Planetárium
Bikiny. Osmdesát let s odhaleným pupíkem

Planetárium

Play Episode Listen Later Jul 10, 2026 3:42


Před osmdesáti lety, v červenci 1946, byly veřejnosti poprvé představeny bikiny. Dvoudílné dámské plavky, na jejichž výrobu není třeba mnoho látky a které odhalují pupíkVšechny díly podcastu Planetárium můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Host Lucie Výborné
Housle, traktor, střelba. Každý film mě naučí něco nového, říká herečka Pavla Gajdošíková

Host Lucie Výborné

Play Episode Listen Later Jul 9, 2026 13:50


Kromě hereckého talentu si do Prahy přinesla hudbu, zpěv, tradiční folklor a zkušenost s hrou na cimbál. Kvůli filmu Tanec s medvědem se ale herečka Pavla Gajdošíková musela naučit hrát na housle. „Měsíc jsem každý den trénovala dvě hodiny. Obdivuju partnera a sousedy, že to se mnou vydrželi,“ říká. Ve snímku Dřevorubec, který představila na festivalu v Karlových Varech, si zase vyzkoušela řízení traktoru i život v prostředí dřevařů.Všechny díly podcastu Host Radiožurnálu můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Večerní Host Radiožurnálu
Housle, traktor, střelba. Každý film mě naučí něco nového, říká herečka Pavla Gajdošíková

Večerní Host Radiožurnálu

Play Episode Listen Later Jul 9, 2026 13:50


Kromě hereckého talentu si do Prahy přinesla hudbu, zpěv, tradiční folklor a zkušenost s hrou na cimbál. Kvůli filmu Tanec s medvědem se ale herečka Pavla Gajdošíková musela naučit hrát na housle. „Měsíc jsem každý den trénovala dvě hodiny. Obdivuju partnera a sousedy, že to se mnou vydrželi,“ říká. Ve snímku Dřevorubec, který představila na festivalu v Karlových Varech, si zase vyzkoušela řízení traktoru i život v prostředí dřevařů.Všechny díly podcastu Host Radiožurnálu můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Play Episode Listen Later Jul 8, 2026 57:55


We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li

Datacenter Technical Deep Dives
Getting Started with Local AI (2/3)

Datacenter Technical Deep Dives

Play Episode Listen Later Jul 6, 2026 48:34


Join us as Du'An digs into the real mechanics of running AI locally and in production - from GPU memory math to multi-agent architectures, observability, and the economics of self-hosted inference. Du'An walks through how model weights and KV cache compete for GPU memory, why continuous batching matters when you have more than a handful of users, and how agent architectures like single-agent, workflow, graph, swarm, and supervisor patterns each solve different problems. You will learn how to instrument your agents with Langfuse for observability and cost tracking, when to use Ollama versus vLLM, how prompt caching can cut provider costs by up to 75%, and why GPUs should never sit idle. Episode two of three - the next episode covers deploying at scale. Timestamps 0:00 Welcome & Introduction 1:47 Du'An's New Role at Akamai Cloud 3:10 Data Privacy and the Case for Self-Hosted AI 7:21 Anthropic and OpenAI as the New Cloud Layer 12:48 Local Models for Specific Use Cases - Cancer Detection Example 15:02 GPU Memory Math - Weights, KV Cache, and Context Windows 19:32 Continuous Batching and GPU Time Slicing 20:03 Observability with Langfuse - Live Demo 27:44 Agent Architectures - Single Agent, Workflow, Graph, Swarm, Supervisor 36:36 Token Economics, Prompt Caching, and GPU Cost Planning 45:32 Ollama vs vLLM - Prototyping vs Production How to find Du'An: https://duanlightfoot.com https://www.linkedin.com/in/duanlightfoot/ Links from the show: https://langfuse.com/ https://github.com/akamai-developers/akamai-workshop-solution-architect-agent https://amzn.to/4bvHn1p https://vllm.ai/

The KE Report
Surge Copper – Breaking Down The Key Metrics and Takeaways From The Berg Project PFS

The KE Report

Play Episode Listen Later Jul 6, 2026 33:37


Leif Nilsson, CEO & Director of Surge Copper (TSX.V:SURG – OTCQB:SRGXF), joins me for a comprehensive update covering the updated Mineral Resource Estimate and Pre-Feasibility Study (PFS), at their flagship copper-molybdenum-silver-gold Berg Project in British Columbia.   Leif mentioned that the completion of the Berg PFS marks an important milestone for Surge and materially advances one of Canada's most significant undeveloped copper projects. Berg now stands out not only for its scale, but also for the quality of its development profile, with long-life production of copper as a primary metal, and industry leading molybdenum and silver output, strong infrastructure advantages, and access to low-carbon hydroelectric power. Just as importantly, this study reflects a great deal of technical work completed since the PEA and provides a more defined foundation for the next stage of advancement, including continued work with First Nations, formal entry into the environmental assessment process, and future feasibility-level studies.   Key highlights from PFS:    Base case after-tax NPV8% of C$4.6 billion, IRR of 24%, and payback period of 2.9 years, based on long-term commodity price assumptions of US$4.75/lb copper, US$20.00/lb molybdenum, US$45/oz silver, and US$3,500/oz gold and an exchange rate of 0.73 US$/C$ At spot prices as of June 2026 (US$6.45/lb copper, US$30.00/lb molybdenum, US$65/oz silver, and US$4,250/oz gold and an exchange rate of 0.73 US$/C$), a spot price sensitivity case generates an after-tax NPV8% of C$9.4 billion, an IRR of 36%, and a payback period of 1.8 years, underscoring the Project's leverage to higher metal prices Maiden Proven & Probable Mineral Reserve of 1.2 billion tonnes grading 0.22% copper, 0.026% molybdenum, 4.1 g/t silver, and 0.02 g/t gold, containing 5.8 billion pounds of copper, 687 million pounds of molybdenum, 160 million ounces of silver, and 0.8 million ounces of gold Updated Mineral Resource Estimate includes Measured and Indicated Mineral Resources of 1.4 billion tonnes grading 0.21% copper, 0.025% molybdenum, 4.0 g/t silver, and 0.02 g/t gold, plus additional Inferred Mineral Resources of 1.0 billion tonnes grading 0.16% copper, 0.027% molybdenum, 4.3 g/t silver, and 0.01 g/t gold 28-year mine life with total production of 8.6 billion pounds (3.9 million tonnes) of copper equivalent (CuEq)1, including 4.9 billion pounds (2.2 million tonnes) of copper, 602 million pounds of molybdenum, and 89 million ounces of silver First 5 years of steady-state production averages 416 million pounds (189 thousand tonnes) of copper equivalent annually, including 270 million pounds (122 thousand tonnes) of copper, 21 million pounds of molybdenum, and 4 million ounces of silver Life of mine average annual production of 308 million pounds (140 thousand tonnes) of copper equivalent, including 176 million pounds (80 thousand tonnes) of copper, 21 million pounds of molybdenum, and 3 million ounces of silver Life of mine C1 co-product cash costs of US$1.95/lb payable CuEq and by-product cash costs of US$(0.17)/lb payable Cu Low life of mine strip ratio of 2.0:1, including waste pre-stripping requirements of 304 million tonnes Initial capital cost of C$4.7 billion and sustaining capital of C$1.7 billion, based on an EPCM execution approach and a three-year construction period, and including a total life of mine contingency of C$715 million, implying initial capital intensity of US$24,416/t CuEq annual average production capacity, and life of mine capital intensity of US$0.55/lb recovered CuEq Selected development case based on a 120,000 tonne per day concentrator and a new 230 kV transmission line connecting the Project to the BC Hydro grid, and downhill overland conveyor transport of ore to the process plant Simple, stand-alone development case based on a single-phase build, conventional open pit mining and processing, with no reliance on phased expansions or third-party major infrastructure     If you have any follow-up questions for Leif regarding Surge Copper, then please email them to me at Shad@kereport.com.   In full disclosure, Shad is a shareholder of Surge Copper at the time of this recording, and may choose to buy or sell shares at any time.   Click here to follow the latest news from Surge Copper   For more market commentary & interview summaries, subscribe to our Substacks:   The KE Report: https://kereport.substack.com/ Shad's resource market commentary: https://excelsiorprosperity.substack.com/     Investment disclaimer: This content is for informational and educational purposes only and does not constitute investment advice, an offer, or a solicitation to buy or sell any security. Investing in equities and commodities involves risk, including the possible loss of principal. Do your own research and consult a licensed financial advisor before making any investment decisions. Guests and hosts may own shares in companies mentioned.    

CzechCrunch Podcast
Ustál pád na dno a prokázal velkou cílevědomost. Síla je v disciplíně, říká brankář Antonín Kinský

CzechCrunch Podcast

Play Episode Listen Later Jul 1, 2026 65:07


Brankář Antonín Kinský je jednou z nejpozoruhodnějších postav tuzemského profesionálního sportu, která v době reprezentačního zmaru ukazuje úspěšnou tvář českého fotbalu. Vždyť tři roky nazpět nastupoval za druholigový Vyškov a teď podepsal nový kontrakt s londýnským Tottenhamem za desítky milionů korun. Ovšem ještě na přelomu jara a zimy byl tamtéž na odpis a pomýšlel na odchod. Jeho příběh je zkrátka plný pádů a vzestupů. „Vždycky jsem měl jasný cíl a disciplínu,“ shrnuje svou cestu třiadvacetiletý gólman v podcastu Money Maker.V rozhovoru se dále dozvíte:

True Story
Kvæleren fra bakkerne 1:5

True Story

Play Episode Listen Later Jun 29, 2026 45:56


I slutningen af 1970'erne rystes Los Angeles af en række brutale fund, hvor unge kvinder findes efterladt på byens øde bjergsider. Efter serien om Rodney Alcala, ”Den charmerende seriemorder”, vender vi nu tilbage til samme tid og sted for at undersøge sagen om Kvæleren fra bakkerne. Fortællingen går helt tæt på en efterforskning, hvor politiet er oppe imod en gerningsmand, der kynisk udnytter falske politiskilte og lovens egen autoritet til at vinde sine ofres tillid. Det er samtidig beretningen om en by i knæ, et komplekst politiarbejde og jagten på en morder, der gemmer sig lige for øjnene af alle.  Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Brain We Are CZ
323: Evoluční Psychologie: Emoce | Mylně přisouzené vzrušení | Kognitivní Kapitulace & Myšlení Rychlé, Pomalé a Umělé

Brain We Are CZ

Play Episode Listen Later Jun 23, 2026 45:05


V tomhle díle jdeme do evoluční a kognitivní psychologie přes konkrétní experimenty, které dodnes nahání husí kůži. Poslušnost autoritě, misatribuce emocí, halo efekt, iluze kompetence i důvěra, která vzniká zvláštní kombinací přátelskosti a autority. Tohle je zkrácená verze (44min) poslouchej celý díl (76min) jen za 100kc/měsíc! Zde: https://open.spotify.com/episode/43sNP0ZiovypG1MP7wSTSD?si=d2aff5d4df564f14Je to epizoda o tom, jak vznikají naše omyly, proč jsou tak přesvědčivé a proč nestačí mít „správná data“, když je mozek nastavený rychle označkovat realitu a jít dál. A taky o tom, že perspektiva není měkké klišé. Je to tvrdý zásah do toho, co cítíš, jak jednáš a jaký svět se ti vůbec ukazuje.Parťáci epizody jsou:KusKakaa - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.kuskakaa.cz⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠  Přináší do Česka čistá ceremoniální kakaa. A proč si takové kakao dopřát? Ukazuje se, že přináší celou řadu benefitů a má velký obsah flavonoidů a polyfenolů. Tak jdi na ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.kuskakaa.cz⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ a zkus jedno z jejich kvalitních kakaí! Doporučujeme to z Kostariky, nebo Peru.N⁠⁠⁠⁠⁠⁠orsan.cz⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Norsan Vyrabí Omega 3 z čerstvého rybího oleje z udržitelného rybolovu nebo z mořských mikrořas. Jdi na Norsan.cz zadej kód bwa10 pro 10% slevu a pořiď si kvalitní OMEGA-3 tvůj mozek a zdraví ti poděkuje.Minutáž:00:00 Úvod05:17 Milgramův experiment a vliv autority07:44 Fenomén šesti stupňů odloučení09:49 Fenomén šesti stupňů odloučení12:06 Hodinář a argument inteligentního stvořitele14:38 Aristoteles, Darwin a počátky evoluční myšlenky20:05 Evoluční důvody sexuální averze k blízkým21:24 Evoluční zpoždění u moderních hrozeb23:05 Model lidské motivace a pocit sounáležitosti25:47 Funkce emocí a sociální cíle27:16 Kognitivní podstata budování důvěry28:34 Misatribuce emocí a experiment na visutém mostě33:14 Valinsův experiment s falešným tlukotem srdce35:21 Falešná interpretace lásky v dramatických vztazích39:12 Pratfall efekt a sympatie z vlastních chyb40:49 Dopaminová křivka a efekt získávání sympatiíPřechod do VIP- Milgramův experiment a síla sociální konformity- Kognitivní kapitulace vůči umělé inteligenci- Nebezpečí slepé důvěry v chyby umělé inteligence- Agentní pareidolie a iluze umělého vědomí- Amnézie zpráv a ztráta kritického myšlení v neznámých tématech- Logická hádanka s pálkou a míčkem- Konformita a vliv skupiny na vnímání barev- Vznik konfliktů a usmíření na letním táboře- Haló efekt a fundamentální atribuční chyba- Matoušův efekt a prohlubování sociálních rozdílů- Kvízový experiment a falešné posuzování kompetence- Efekt blábolení a jeho využití v politice- Síla přerámování vnímání reality a vědomé změny prožitku- Buddhistický příběh o voru a opouštění zátěže- Indické podobenství o slepcích a slonovi- Závěrečný citát o inteligenci a lidském charakteru

Verkligheten i P3
Felix sa ifrån mot rasism på nattbussen – blev misshandlad

Verkligheten i P3

Play Episode Listen Later Jun 23, 2026 26:12


Två dagar efter får Felix syn på två av de män som slog ner honom. Nu står han blåslagen framför dem. Vad ska han göra? Konfrontera eller förlåta? Lyssna på alla avsnitt i Sveriges Radios app. Reporter & ljuddesign: Leontine OlsbjörkProducent: Gustav AsplundSlutmix: Astrid AnkarcronaVerkligheten görs av produktionsbolaget Filt.UTSKRIFT AV AVSNITTET:– Och jag kände först en ångest av att gå ut för att jag. Ja, jag har precis varit med om någonting väldigt traumatiskt och jag ser ut som jag gör. Men tänker väl ändå att det här ska jag ska inte skämmas eller jag ska väl inte behöva stanna inne för att jag har varit med om det här.Felix är blåslagen i ansiktet. Två dagar tidigare har han blivit misshandlad på en nattbuss.När han sa ifrån mot några unga män som uttryckte sig rasistiskt.– Jag står där och värmer mig vid en eld och ser då två av de här. Killarna kommer gåendes.jag säger då till min familj att där är de ju.Nu står han mitt emot två av dom unga män som misshandlade honom. Nu, ska han konfrontera killarna – och förlåta dem.***Första gången jag ser Felix är i TV4 Nyhetsmorgon. Jag står i green room och pillar nervöst på en kaffekopp. Om en liten stund ska jag intervjuas i direktsändning.Du kanske har hört min berättelse. För Verkligheten har jag berättat om när jag tog ett DNA-test och hittade över 30 halvsyskon. Och den här morgonen ska jag prata om det i Nyhetsmorgon. Det är december och i sminket pratas det om det stundande luciatåget. Framför mig sitter flera tv-skärmar uppsatta på väggen. I dem ser jag Felix.Han sitter i en soffa inne i studion. Drygt två veckor tidigare har han blivit misshandlad på en nattbuss, efter att ha sagt ifrån mot några unga män som uttryckt sig rasistiskt. Jag har fått pushnotiser om misshandeln tidigare, men nu ser jag Felix berätta om kvällen. Fortfarande med blåmärken under ögonen. Och det är en sak jag fastnar för. För Felix säger att han har förlåtit dem.Intervjun tar slut och innan Felix försvinner vidare genom korridoren och jag in i studionhinner vi hälsa snabbt. Men det är något som jag inte riktigt kan släppa.Några månader senare sitter jag på ett tidigt morgontåg norrut, mot Vännäs i Västerbotten. Jag ska träffa Felix, igen. Felix möter mig upp vid tågstationen– Här har vi Myranparken.Vi går genom den lilla bruksorten några mil utanför Umeå.– Min farfar bodde i det här gula huset där. Då såg parken lite annorlunda ut när vi var små.Felix och hans syskon brukade leka här som barn.– Det var ganska mycket mer växtlighet här, så man kunde klättra i några träd-ish buskar och sådär. Och så har man ett starkt minne av en rutschkana som var lite farlig, men det är lite andra regler idag.Han växer upp i en annan liten ort, Rundvik, men han var ofta i Vännäs som barn. Och som vuxen flyttar han hit.– Pappa var väl inget stort fan av Vännäs. Han sa till mig innan han gick bort till 2020. Han sa att du får fan inte flytta till Vännäs. Men då träffade jag en Vännäs-bo. Så han får vända sig graven om han vill.När Felix är liten separerar hans föräldrar. Hans pappa missbrukar och kontakten är sporadisk.– Han var också en klok man. Var också närvarande när vi väl var där. Mycket starka minnen om att vi hittade på saker. Så man såg upp det handeln då. Fick mycket från honom i just hur jag är som person. Väldigt, väldigt... Jag analyserar mycket, jag tänker, reflekterar. Och det har inte alltid heller varit lätt.Felix mamma är från Brasilien och familjen är den första med invandrarbakgrund i byn.– Man fick höra mycket när man växte upp. Väldigt mycket glåpord. Och i och med att min familj inte jobbade heller och skaffade många barn så var man ganska utsatt på grund av det. Det var många som visste vilka vi var utan att man hade pratat med dem. Och mycket fördomar.De har det tufft ekonomiskt. Felix och hans syskon går ofta klädda i ärvda kläder och det händer att dom går i fel kläder för säsong. De sticker ut.– ´Det var en inre stress hela tiden. Man skämdes för att vi hade så mycket syskon som levde om och inte kunde sköta sig. Enligt en själv då. Idag var ni ju barn så det är svårt att klandra dem. Men det var så man tänkte då i alla fall.På skolgården fortsätter fortsätter glåporden.– Men det var ju mycket på skolan. Man fick ju höra diverse. Allt från svartskalle till turkjävel.Ibland säger han ifrån och andra gånger skrattar han bara med– Men många gånger svalde man det också. Det var som att bita ihop och hantera det. Jag har svårt att använda ordet rasism. Jag vet inte varför. Men det var helt klart rasistiskt och det var helt klart okunskap.Felix känner sig utanför. I ett försök att passa in bleker han sitt mörkbruna hår.– Det slutade inte så bra för att det blev orange. Det var ytterligare en sak som var såsynlig att man gjorde ett försök som inte slutade bra.Hemma får Felix och hans syskon höra att det är viktigt att de sköter sig i skolan och bland vuxna.– Jag var väldigt mån om att följa regler och hjälpa till där det behövdes städa undan och respektera de äldre. Upplevelserna när jag var liten, det har ju absolut format mig idag. Jag har noll tolerans emot mobbning generellMen skolan var också en plats där Felix blev sedd och även om det stundtals var tufft hemma,så kommer han från en kärleksfull familj. Vi spolar fram till en fredagskväll i november förra året, då utanförskapet och hatet gör sig påmint igen. För Felix växer upp. Han är social och utåtriktad. Han brinner för barn och unga, jobbar på fritidsgård och som ungdomstränare. Han börjar plugga till lärare med hoppar av och börjar jobba inom skogsindustrin.– Jag hade precis varit på en AW.På en arkadhall i Umeå står Felix och hans kollegor bland blinkande pinball-maskiner och gamla retrospel.– Så jag hade bokat det och så hade jag fixat pizzor till alla och så där. Så vi hade varit där hela kvällen och umgåtts och haft jättekul.Det är en trevlig kväll. Men så börjar samtalen förändras.– En sak som som kom upp i samtalet där under kvällen var just politik. Där jag blev lite förvånad över folks åsikter och värderingar och så där. Men vi valde sedan att gå vidare då.AW:n fortsätter. Felix är på gott humör när han skiljs från sina kollegor. Han ska ta nattbussenhem till Vännäs och beger sig till busshållplatsen.– Jag hade börjat prata med två med två kvinnor där på busshållplatsen och skrattat och skämtat och haft kul. Men det där låg lite grann i bakhuvudet också, just över den politiska den diskussionen vi hade Det var väl därför jag kanske också initierade i det här samtalet med de här kvinnorna.Kvällens samtal gnager fortfarande i honom när han börjar prata med kvinnorna. De går på bussenoch sätter sig långt fram där fortsätter samtalen.– De sätter sig ganska nära mig så att jag vänder mig och börjar prata med de här två kvinnorna. Och ja, vi pratar om allt möjligt och jag frågar hur länge de har bott här i Vännäs och hur de ser på på klimatet här i kommunen. Vi delar lite åsikter och tankar och sådär då. Och vi sitter och pratar egentligen hela vägen från från Umeå till Vännäs by där de hoppar av.Men när kvinnorna kliver av bussen förändras stämningen snabbt Felix hör några killar längre bak i bussen prata han hör att de säger något som slår an i honom. De säger n-ordet flera gånger– Så jag går bak och sätter mig och sätter mig och sätter mig och sätter mig framför dem och hälsar och säger såhär till grabbar. Så här kan ni inte uttrycka er. Ni måste tänka er för på hur ni pratar och hur ni uttrycker det, för det här är inte okej. Jag blir väl bemött ganska ignorant, de tycker att de får säga vad de vill och de har inte sagt någonting som inte är lämpligt.Några säten längre fram sitter en kvinna och lyssnar. Hon vänder sig om.– Och också säger till den här kvinnan är ju då svart. Så hon har ju slutat höra på allt det här vad de har sagt, så hon vänder sig om också och säger, säger åt dem. En av de här personerna blir då ganska, störig. Han tycker, han försöker liksom provocera mig och så där. Och så efter lite, lite fram och tillbaka så är vi då framme i Vännäs näst sista hållplatsen och då säger de ”ska vi kliva av här”?Varför jag svarar ja, det kan vi göra.Felix reser sig upp och går mot bussens utgång.– Men då upptäcker jag att jag har glömt min väska på mitt säte.Han vänder tillbaka in genom bussen– Var på den här personen skriker “vad fan gör du” och tar stryptag på mig. Och jag ramlar som över honom där han satt och då slår han mig två gånger. Felix blir slagen med knytnävsslag i ansiktet. Han försöker hålla fast killen för att avvärja slagen och skydda sig. – Och jag säger till honom att att är det bra nu? Var på det någon bakifrån börjar slå mig. Och då tar jag tag i honom också och håller fast de två. Och så började en tredje slå mig. Ingen säger någonting. Jag blir slagen, slagen, slagen.Killarna fortsätter slå Felix medan bussen åker vidare mot ändhållplatsen– När bussen stannar så släpper jag sista personen som jag håller i. Killarna försvinner iväg. Felix halvligger över sätena längst bak i bussen. Ingen på bussen hjälper Felix. Och jag sätter mig upp och känner hur det rinner blod från ansiktet. Så till slut så ställer jag mig upp och tittar upp och då ser jag busschauffören stå där och han säger ingenting. Så jag kliver av bussen och bussen kör iväg.Skärrad och blodig står Felix på busshållplatsen mitt i natten. Han plockar upp sin mobil och slår numret till 112. Efter en stund kommer både polis och ambulans, polisen ställer frågor och ambulanspersonalen undersöker Felixs ansikte och hals. Han förs till akuten i Umeå. Samtidigt försöker Felix på tag på sin sambo, men hon svarar inte. – Så jag skickade en bild i en familjegrupp vi har. Och förklarade vad som hade hänt lite snabbt. och att de måste försöka få tag på henne. På bilden är hans högra öga igensvullet. Ansiktet är blodigt och rödflammigt.– Jag grät och jag fick panik. Hur ska jag berätta det här för min sambo? Mina barn ska se mig också.Felix har fått flera frakturer i ansiktet, i kindbenen och näsan. Först framåt morgonen får han tag på sin sambo.– Jag kommer ihåg att jag sa till henne att jag var på en säker plats och ändå mådde bra. Det gör det jobbigt nu också när jag säger det högt att jag inte kommer ihåg hur hon reagera.samtidigt går tankarna på högvarv.– Jag var också så chockad över att det var ett sådant våldskapital. Det är inte första gången jag sitter på den här bussen hem efter en utekväll och har haft samtal med diverse personer och stoppat bråk mellan andra personer. Jag brukar kunna se var saker och ting är på väg. Här hade jag ingen aning. Jag förstod mig inte alls. Jag kunde inte se det framför mig. Jag kände ingenting som att det skulle bli sådant här våld. och så tyckte jag synd om dem för just av den anledning, att det va var verkligen så oväntat så det gjorde ont i mig att unga människor är så arga.Det har gått två dagar sedan misshandeln. Det är skyltsöndag. Byn är full av människor. Det är ponnyridning, marknad och tomtebesök.– Jag kände först en ångest av att gå ut. Jag har precis varit med om något väldigt traumatiskt och jag ser ut som jag gör. Jag tänker att jag inte ska skämmas för att jag har varit med om det här.Trots att han är synligt blåslagen i ansiktet följer Felix med sin familj till byns torg.– Jag står där och värmer mig vid en eld och står och pratar med min mamma och min syster. Då ser jag två av de här killarna komma gående. De ser att jag ser dem och då vänder de ganska fort. Jag säger till min familj att där är de. Mamma säger åt mig att inte följa efter.Men Felix följer efter dem. De skyndar iväg till en parkerad bil.– De hinner sätta sig i bilen men jag går fram och knackar på förarfönstret. Jag sätter mig på huk och hälsar och säger hej känner ni igen mig Killarna nekar. Och Felix blir allt mer irriterad.– Så ni känner inte igen mig? Från bussen i fredags. Då säger de att, nej, vilken buss? Det här är samma attityd ni hade då som nu. Det är inte okej. Okej, ja, men hur ska vi lösa det här nu då? Jag säger något i stil att de ska se på mig och titta vad de har gjort med mig. Då mjuknar de lite och ber dem ursäkt för det de gjorde. Då säger jag till dem att de gärna får be sina kompisar ringa mig också. Jag lämnar mitt nummer och mitt namn. Sen kör de därifrån.Men Felix känner sig inte klar. Han söker på bilens registreringsnummer. han får fram adressen, men hejdar sig.– Jag kände att jag inte kan träffa honom eller dem ensamma. Det kändes inte bra. Istället tar han reda på var hans föräldrar bor.– Jag åker dit och parkerar bilen. Jag tänker nog inte så mycket på vad jag ska säga annat än att presentera mig och berätta vad som har hänt. Jag knackar på dörren och så öppnar en man.Det visar sig att mannen som öppnar är styvfar till en av killarna i bilen– Jag hälsar och säger hej. Jag heter Felix. Jag berättar att jag har blivit misshandlad på bussen denna fredags. Jag frågar om jag får komma in. Han bjuder in mig och vi sätter oss i deras vardagsrum.Det får han. Styvpappan bjuder felix på en kopp kaffe, och de sätter sig ned i vardagsrummetoch pratar.– Och jag säger väl att jag inte vet hur jag ska hantera det här på något annat sätt än vad jag gör just nu. Jag frågar om det var konstigt att jag var där och han tyckte att det inte var det. Vilket jag har varit tacksam för.En stund senare kommer mamman till en av killarna hem.– Hon frågar mig då om hon ska ringa hem grabbarna igen då. Så de kommer dit också.Nu möter Felix två av killarna igen. De sätter sig ned och pratar.– Jag ville visa att så här ser jag ut idag. Jag åkte dit av känslan att. Jag tänker inte gömma mig för det här. Utan personerna måste verkligen få se vad de har gjort och vad... Det finns en person på andra sidan här också. Och att saker får konsekvenser.Killarna ber om ursäkt igen och den här gången svarar han att han förlåter dem– Jag känner väl en sorg över att, att de valde det här, för det är också någonting de gjorde.Men att de inte har haft samma möjligheter kanske eller samma förutsättningar för att kunna reglera sig själv eller hantera sig själv. Eller fått stöd och hjälp i just känslor.– För jag har också varit i den åldern och jag var också ganska arg. Men jag var inte arg på... Men jag var mer arg på var jag kommer ifrån. Jag kommer själv från förhållanden som... Som har varit långt ifrån optimala. Jag kommer från ett kärleksfullt hem, absolut.Men jag kommer ifrån mycket och stora utmaningar.Felixs val att förlåta killarna tas emot med blandande känslor och en vän till honom är orolig att han inte kommer medverka i polisutredningen.– Och han har också sagt till mig rakt upp och ner. Att... Varför ska du vara så här för? Varför ska du... Varför ska du ens förlåta dem? Du är inte bättre än oss andra. Han har sagt. Och det är jag inte. Inte alls. Men sen kanske man också måste tillåta sig själv att få vara. Arg då eller ledsen eller sårad eller utsatt.Och såhär ett halvår senare märker han att misshandeln, precis som uppväxten, har satt sina spår– Och jag har ju fått frågan tidigare om jag kommer fortsätta ingripa och säga ifrån. Och jag har väldigt tryggt sagt ja, det här kommer inte förändra mig. Och ja, det har det gjort. Jag är mer osäker nu än tidigare. Att ha allt det där i ryggsäcken och så får man växa upp i en mindre ort med mycket fördomar och rasism. Det skapar ett annat förhållningssätt till livet än vad många här i Sverige kanske förstår.– Jag är supertacksam för det Sverige jag har växt upp i. Hela skyddsnätet som vi har haft är därför jag sitter här idag. Jag hade kunnat vara på en helt annan plats om jag inte hade haft det. Det måste vi värna om, att vi tar hand om våra medmänniskor som inte har allt på det torra.Efter misshandeln samlas människor i Vännäs i en manifestation mot våld och rasism. När de tågar genom byn hörs ropen: “Alla hjärtan har samma färg.”– All den kärlek och värme jag fick från hela Sverige, från olika människor. Det var personer som skrev brev till mig och personer som hörde av sig på alla möjliga olika medier. Och det visar ju på att det finns en kraft och att det finns människor där ute som på riktigt bryr sig. Och det är lätt man glömmer det, tänker jag. Det är lätt att man glömmer bort att det här berör så många. Det har jag inte tänkt på länge nu så det kanske jag ska försöka tänka på oftare. Just att den värmen och kärlek jag fick efter allt det här. Jag ska med mig det och att det kanske också kan få andra att agera i framtiden också. Om de ser eller hör eller upplever någonting.

Kvällspasset i P4
Kvällspassets midsommarspecial!

Kvällspasset i P4

Play Episode Listen Later Jun 19, 2026 43:02


Kvällspasset är på plats hemma hos Karl Fredrik Gustafsson på Österlen för att prata midsommarminnen, blommor, mat och musik! Lyssna på alla avsnitt i Sveriges Radios app. Ett nyfiket och underhållande aktualitetsprogram med lyssnaren i fokus.Tillsammans med floristen och tv-profilen Karl Fredrik, artisten Diane Emerita, matinfluencern Julia Tuvesson och lyssnarna tar vi oss an midsommarafton.Julia bjuder på midsommarpizza, Karl Fredrik fixar kransarna och Diane sjunger bland annat en vacker tolkning av Lasse Berghagens ”En kväll juni”.

The Right Idea
Texas' 765kV Transmission Lines EXPOSED: Permian Basin Crisis, ERCOT Failures & Why $33 Billion Is a Terrible Idea | The Right Idea #112

The Right Idea

Play Episode Listen Later Jun 18, 2026 23:56


In this episode of The Right Idea, Derek Cohen sits down with Carson Clayton, TPPF Life Powered Campaign Director, to break down the controversial 765 KV transmission lines (Strategic Transmission Expansion Plan / STEP).Texas lawmakers and activists are sounding the alarm over a $33 billion plan to build massive ultra-high voltage transmission lines across the state — cutting through pristine Hill Country and farmland — instead of addressing the root cause: Texas' broken energy market that over-subsidizes intermittent wind and solar while under-building reliable dispatchable power.Featuring powerful clips from Senator Kevin Sparks and Representative Brad Buckley.Key Topics:Why the Permian Basin — one of the most energy-rich regions in the world — needs power imported from Central TexasHow federal subsidies, ESG pressure, and ERCOT's energy-only market are distorting investmentThe massive cost to ratepayers and landownersWhat real market reform looks like to prevent blackouts and unnecessary transmission boondogglesTimestamps:00:00 - Welcome & Introduction to the 765 Lines Controversy01:23 - What Are the 765 KV Transmission Lines?02:41 - Senator Kevin Sparks on the Permian Basin Plan04:40 - Why Wind & Solar Boom Created a Reliability Crisis08:37 - Senator Sparks on ERCOT Market Failures & Subsidies09:39 - How the Energy-Only Market Rewards Unreliable Power12:19 - Federal Policy, ESG, and Renewable Credits Driving the Problem15:12 - Rep. Brad Buckley: Pause the Project & Reform the Market17:02 - Proposed Market Reforms (SB 715 & Reliability Standards)20:01 - The Coming Reliability Cliff & Why Transmission Is Just a Band-Aid21:22 - How Much Gas Generation Would Make the 765 Lines Unnecessary?If you care about Texas energy independence, affordable electricity, property rights, and keeping the lights on, this is a must-watch.

444
Borízű hang # 274: Hallgatóbüdösségi és bongépítő verseny, valamint a „Hülye vagy kretén?” szellemi vetélkedő országos döntője a Fidesz-frakcióban [rövid verzió]

444

Play Episode Listen Later Jun 14, 2026 50:35


Az előfizetők (de csak a Belső kör és Közösség csomagok tulajdonosai!) már szombat hajnalban hozzájutnak legfrissebb epizódunk teljes verziójához. A hétfőn publikált, ingyen meghallgatható verzió tíz perccel rövidebb. Itt írtunk arról, hogy tudod meghallgatni a teljes adást. Radikális fordulat dezodorügyben. NBA-döntő vs. vébé. Ellen-Trumpok halálfej-tetoválással. Tajtékpipa és gravity bong. Németh Balázs elmeállapota. Mikor csuknak már le valakit? Az osztrákoknál az alma is jobb. 00:00 Tartalomjegyzék. A 96 órás dezodor. Néhány támogatói jegy még van a mosdatlan Borízű Live-ra. 04:07 Személyiségi jogi védelem a Budaörsi uszodában. Nincs is semmi baj a dezodorokkal.08:52 Mérsékelt vébéhőemelkedés. A jó 94-s amerikai vébé. Diana Ross tizenegyese. Orbán Viktor megnézi az Üzbegisztán-Kolumbiát. 14:56 NBA-világbajnokság-együttállás. Wemby olvasgat. Spike Lee: Knicks in six. A csapatépítő Gregg Popovich. 21:19 A kosár, amivel a negyedik meccset megfordította a Knicks. Halt time show a Wu-Tang Clannel. Shaq bulizik a stúdióban. Amerikai sportrajongók politikai hovatartozása. 26:29 Platner, Talarico, Newsom és a demokraták laboratóriumban megalkotott ellen-Trumpjai. Mamdani Arsenal-drukker imája. Mamdani végigtippeli a vébét. 31:39 Kvíz: tajtékpipa. Te miből szívtál cracket? Pipázás a rendszerváltás körül. 36:33 Első 444 Bongépítő Verseny. Vödrözni egészséges! Bottle bag és egyéb gravity bongok. Cindy Breakspeare és Pascaline Bongo. 41:40 A Tisza-négyötöd robotosai. Ha jól megy a gazdaság, nem csukják le Orbánt. 43:04 Uj Péter Ausztriában. A Billában bezzeg rendes alma van! Cosmic crisp és Jack Herer. Az almabong óvodás szint. See omnystudio.com/listener for privacy information.

Alexa's Input (AI)
How vLLM and llm-d Changed AI Inference with Rob Shaw

Alexa's Input (AI)

Play Episode Listen Later Jun 3, 2026 102:59


In this episode of Alexa's Input (AI), I sat down with Rob Shaw from Red Hat to talk about how AI inference evolved from a simple model serving problem into a large-scale distributed systems problem.We explored the infrastructure shifts behind modern LLM serving, including how vLLM and PagedAttention changed the economics and efficiency of inference, why KV cache management became one of the most important bottlenecks in production AI systems, and how orchestration layers like llm-d are emerging to coordinate distributed inference.We also discuss:how LLM inference differs from traditional model serving runtimesKV cache, prefix caching, and cache-aware routingwhy throughput and latency became major infrastructure challengeslong-context agents and repeated inference callsdistributed inference on Kubernetesintelligent routing, flow control, and load balancingprefill/decode disaggregationenterprise AI deployment realitiesvLLM has become one of the most important open-source projects in AI infrastructure, and llm-d represents a newer shift toward treating inference as a coordinated distributed system rather than just a single runtime problem.If you want to better understand the systems layer beneath modern AI applications, this episode is a deep dive into where inference infrastructure is heading next.General Podcast LinksWatch: ⁠⁠⁠⁠⁠⁠https://www.youtube.com/@alexa_griffith⁠⁠⁠⁠⁠⁠Read: ⁠⁠⁠⁠⁠⁠⁠⁠https://alexasinput.substack.com/⁠⁠⁠⁠⁠⁠⁠⁠Listen:⁠⁠ ⁠⁠https://creators.spotify.com/pod/profile/alexagriffith/⁠⁠⁠⁠More: ⁠⁠⁠⁠⁠⁠https://linktr.ee/alexagriffith⁠⁠⁠⁠⁠⁠Learn more about the host atWebsite: ⁠⁠⁠⁠⁠⁠https://alexagriffith.com/⁠⁠⁠⁠⁠⁠LinkedIn: ⁠⁠⁠⁠⁠⁠https://www.linkedin.com/in/alexa-griffith/⁠⁠⁠⁠⁠⁠Find out more about the guest at:LinkedIn: https://www.linkedin.com/in/robert-shaw-1a01399a/ Red Hat Articles: https://developers.redhat.com/author/robert-shawGithub: https://github.com/robertgshaw2-redhat ResourcesvLLM Website: https://vllm.ai/vLLM GitHub Repository: https://github.com/vllm-project/vllmllm-d Website: https://llm-d.ai/llm-d GitHub Repository - https://github.com/llm-d/llm-d KeywordsAI inference, VLLM, LMD, distributed inference, GPU optimization, open source AI, Kubernetes, multi-cluster deployment, AI infrastructure, enterprise AI AI infrastructure, Kubernetes, model optimization, speculative decoding, mixture of experts, AI deployment, performance tuning, AI systems, neural network scaling Key TopicsEvolution of vLLM and llm-dDistributed inference and routingGPU utilization and performance optimizationOpen source AI infrastructureEnterprise deployment challenges and solutions Standardization in Kubernetes for NIC exposurePerformance optimizations: quantization and speculative decodingMixture of experts architecture and parallelism strategiesFlow control and request scheduling in AI systemsEmerging hardware for AI inference, Cerebras processorReinforcement learning and AI system supportModular architecture of vLLM and ecosystem projects

This Week in Machine Learning & Artificial Intelligence (AI) Podcast
How to Engineer AI Inference Systems with Philip Kiely - #766

This Week in Machine Learning & Artificial Intelligence (AI) Podcast

Play Episode Listen Later Apr 30, 2026 54:51


In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed systems, and where the line sits between inference and model serving. Philip shares how research-to-production can move in hours, not months, and why understanding “the knobs” of inference—batching, quantization, speculation, and KV cache reuse—lets teams design better products and SLAs. We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms, discuss GPU lifecycles, and survey today's runtime landscape, including vLLM, SGLang, and TensorRT LLM. Finally, we look ahead to agents and multimodality, making the case for specialized, workload-specific runtimes when performance and efficiency matter most. The complete show notes for this episode can be found at https://twimlai.com/go/766.