Podcasts about asic

  • 675PODCASTS
  • 2,926EPISODES
  • 34mAVG DURATION
  • 5WEEKLY NEW EPISODES
  • Aug 17, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about asic

Show all podcasts related to asic

Latest podcast episodes about asic

The Quicky
Your AM Headlines + Will A Gun Buyback Make Us Safer?

The Quicky

Play Episode Listen Later Aug 17, 2026 11:21 Transcription Available


HEADLINES: NSW households could be shielded from rising power and water costs linked to the state's data centre boom, under a new government framework announced yesterday. A deepfake of Anthony Albanese has been used to dupe Australians out of seven point four million dollars in fake investment scams, ASIC is warning. Tributes are flowing for American actor Hayden Panattiere who died at the age of 36. Australian health experts fear vaccine scepticism from the US is spreading. A 13-year-old Michigan girl, shot four times during a rampage that killed six, has ben hailed a hero after still managing to call 911 and identify the gunman. And one of Australia's favourite spreads is about to get a whole lot colder. If this brings anything up for you, Lifeline is on 13 11 14, and Beyond Blue is on 1300 22 46 36.GET IN TOUCHGot a story, news tip-off, feedback or dilemma?Send us a voice note or email us at thequicky@mamamia.com.au HELPFUL LINKS: Become a Mamamia subscriber and get an all-access pass to everything we make, including exclusive podcasts and early listening, subscriber-only articles, monthly giveaways and our home workout app, MOVE. Get access to Very Peri, Mamamia's exclusive perimenopause series, for just $59. 25 world-leading experts, over 20 on-demand sessions, available now. We’ve sorted through the noise so you don't have to. Go to veryperi.com.au today. You hot? Same. CREDITSHost: Charlotte Mortlock Audio Producer: Scott Stronach Group Executive Producers: Georgie Page and Tamsin Rose Check out The Quicky Instagram here and our TikTok here Discover more Mamamia podcasts here Did you know some of our shows are now in video on the Apple Podcast app? Make sure your phone is up to date and check it out here! Mamamia acknowledges the traditional owners of the land on which we have recorded this podcast.Become a Mamamia subscriber: https://www.mamamia.com.au/subscribeSee omnystudio.com/listener for privacy information.

Hard Reset
E98 - Riddles (Prof. Oded Margalit)

Hard Reset

Play Episode Listen Later Aug 17, 2026 61:36


מה זאת חידה?לא, באמת, תנסו לתת הגדרה.בעיה שצריך לפתור? גם תרגיל במתמטיקה הוא בעיה שצריך לפתור.שאלה עם טוויסט? לא לכל חידה יש טוויסט בתשובה.בעיה שדורשת חשיבה? כלומר Debugging הוא חידה?אז איפה בדיוק עובר הגבול בין בעיה לחידה?כדי לפתור את החידה הזאת הבאנו מישהו שאוהב חידות בצורה קצת מוגזמת: פרופ' עודד מרגלית.עודד הוא פרופסור למדעי המחשב באוניברסיטת בן גוריון, המדען הראשי של NextSilicon ויו"ר עמותת CodeGuru. הוא פותר חידות מילדות, טס לכנסי חידות בינלאומיים, מלמד קורס אקדמי על חידות, ולפי מקורות זרים - אפילו מבשל בחידות.על מה דיברנו בפרק?מה זאת בכלל חידה? ומה ההבדל בין חידה, תרגיל, בעיה ושאלה?אילו סוגי חידות קיימים?איך ניגשים לפתור חידה?מה הופך חידה לחידה טובה?למה שואלים חידות בראיונות עבודה?האם חידות באמת מודדות יכולת פתרון בעיות?וגם, אי אפשר לדבר שעה על חידות בלי לנסות לפתור כמה:כל אחד מאיתנו הגיע עם חידה במטרה לנסות ולהכשיל את האחרים. אחרי שהאזנתם לפרק, ופתרתם את כל החידות בו, מוזמנים להצטרף לקהילה שלנו >>>https://hardreset.co.il/communityונשמח לשמוע בתגובות:מה החידה הכי אהובה עליכם? פרק 98 - RiddlesHard Reset - הפודקאסט של קהילת Hardware Engineering Israel.מוזמנים ליצור איתנו קשר במייל podcasthardreset@gmail.com כמו שהבטחנו בפרק - הנה כמה קישורים למי שרוצה לצלול עמוק יותר לעולם החידות:עודד במגזין מדעי המחשב של הטכניון - עמ' 32–33, גיליון 9https://web.archive.org/web/20191124132318/http://www.cs.technion.ac.il/magazine/09/homepage09.pdfאליפות הבידוד עם קשתhttps://www.youtube.com/playlist?list=PL4WTxrn1Fn6DOqrTqSDvxnjPtz6R22j_U"זרוק מספר" - החידות של עודד בעיתון הארץhttps://www.haaretz.co.il/riddles/numbersסרטון גיוס AIhttps://www.linkedin.com/feed/update/urn:li:activity:7269733034445189121החידה השימושית - חידות וראיונות עבודהhttps://www.linkedin.com/posts/riddles-hiring-nextsilicon-ugcPost-7222239392612909057-YR7Cההרצאה של עודד בכנס חידות - קורס "הוכחות שגויות"https://www.youtube.com/watch?v=CYKXBkGkpUI

Proactive - Interviews for investors
Hive Digital Technologies targets $1bn revenue with new $350m AI data centre deal

Proactive - Interviews for investors

Play Episode Listen Later Aug 17, 2026 5:58


HIVE Digital Technologies Ltd (TSX:HIVE, NASDAQ:HIVE, FRA:YO0, BVC:HIVECO) executive chairman Frank Holmes told Proactive's Stephen Gunnion that a new five-year agreement has taken Buzz HPC's contracted active ARR to around $180 million, and represents a key milestone in the company's strategy to build $300 million in revenue from tier-one data centres and Bitcoin mining, alongside a separate $300 million AI Gigafactory target. Holmes said the collective long-term goal is to reach one billion dollars in revenue, with a particular focus on the AI business. The new contract relates to a facility in Merritt, British Columbia, three hours from Vancouver, powered by hydroelectricity and backed by a strategic partnership with Bell Canada. Holmes said Hive is deploying approximately $185 million in the cluster and expects the 2,016 Blackwell Ultra GPUs to be operational in the fourth quarter of this year, with the run rate expected to ratchet up from over $3 million per month currently toward $20-$30 million per month, and ultimately $50 million per month. Holmes drew a sharp contrast between GPU and ASIC economics, noting that older GPU chips earn around $1 per hour per chip versus $0.10 for an ASIC chip, while the latest NVIDIA chips can push toward $4 per hour. He also pointed to GPU longevity - potentially seven years of value - compared with ASIC chips, which he said are largely worthless after four years, as a key reason for confidence in long-term returns. On financing the build-out, Holmes noted that large alternative capital providers including Blackrock and Blackstone are among those lending against GPU assets, reducing reliance on traditional bank financing. Hive has previously outlined plans to deploy more than 120,000 GPUs over the next two years through its Buzz HPC division. Read Proactive's Editorial Policy here: https://www.proactiveinvestors.co.uk/pages/editorialPolicy #HiveDigital #HIVE #BuzzHPC #AIdatacentre #GPUcomputing #Bitcoin #NVIDIA #Canada #techinvesting #growthinvesting

Investing Experts
What's Up With Tech?

Investing Experts

Play Episode Listen Later Aug 13, 2026 33:27


Sara Awad from Tech Contrarians talks tech's tough year (0:40) Great results being viewed as not good enough (3:20) Semiconductors have more downside, but potential remains (6:50) Memory dynamics (8:40) Thinking about 2027 (14:50) ASIC shift favors ARM (19:45) Are we in a bubble? (23:30)Show Notes:Is The Market Wrong On SpaceX? TheTechTalk Podcast Ep. 1Fundamentals Over EverythingThe Cure For FOMO With Tech ContrariansEpisode transcriptsFor full access to analyst ratings, stock quant scores and dividend grades, subscribe to Seeking Alpha Premium at seekingalpha.com/subscriptions

SBS Malayalam - എസ് ബി എസ് മലയാളം പോഡ്കാസ്റ്റ്
കാർ ഇൻഷുറൻസ് പ്രീമിയത്തിൽ കുതിപ്പ്; കമ്പനികൾക്ക് മുന്നറിയിപ്പുമായി ASIC

SBS Malayalam - എസ് ബി എസ് മലയാളം പോഡ്കാസ്റ്റ്

Play Episode Listen Later Aug 13, 2026 5:33


ഓസ്‌ട്രേലിയയിൽ കാർ ഇൻഷുറൻസ് പ്രീമിയം പണപ്പെരുപ്പത്തേക്കാൾ വേഗത്തിൽ ഉയരുന്നതായി ഓസ്‌ട്രേലിയൻ സെക്യൂരിറ്റീസ് ആൻഡ് ഇൻവെസ്റ്റ്‌മെന്റ് കമ്മീഷൻ റിപ്പോർട്ട്. കഴിഞ്ഞ ഒരു വർഷത്തിനിടെ പ്രീമിയത്തിൽ എട്ട് ശതമാനം വർദ്ധനവുണ്ടായതായി റിപ്പോർട്ട് വ്യക്തമാക്കുന്നു. വാർത്തയുടെ വിശദാംശങ്ങൾ കേൾക്കാം മുകളിലെ പ്ലെയറിൽ നിന്നും...

Lawyers Weekly Podcast Network
Governance lessons from ASIC v Bekier & Ors

Lawyers Weekly Podcast Network

Play Episode Listen Later Aug 12, 2026 21:38


The recently concluded proceedings brought against senior executives of the Star Entertainment Group offer pertinent takeaways for boards and general counsel alike. In this episode of The Lawyers Weekly Show, host Jerome Doraisamy welcomes back Corrs Chambers Westgarth head of investigations and inquiries Abigail Gill to discuss the state of affairs for investigations work, how the matter of ASIC v Bekier & Ors came about and what the Federal Court held, the headline lessons for boards and GCs and their teams, what those takeaways mean for FY2026–27 and beyond, and implications for AI use moving forward.

ai lessons governance federal court asic gcs star entertainment group corrs chambers westgarth jerome doraisamy
The Lunar Society
Ryan Greenblatt – What happens once AI can automate AI research?

The Lunar Society

Play Episode Listen Later Aug 11, 2026 132:32


Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in Large Language Models", and is currently working on a third party investigation into the OpenAI/HuggingFace incident. In my opinion, he's one of the most interesting thinkers on the future of AI.Had him on to discuss/debate recursive self-improvement. This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.I've historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan's median for when we automate AI R&D is 2031.We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what's happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.The first piece of advice you get when you're learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!Watch on YouTube; read the transcript.Sponsors* Antithesis is a software testing platform that finds the failures no human or AI could ever anticipate. It runs thousands of copies of your code inside a fully deterministic computer, injecting faults and steering each trajectory toward the most insidious bugs. This lets you find critical issues in minutes rather than waiting months for your users to uncover them. Learn more at antithesis.com/dwarkesh* Jane Street's back with a new puzzle. They designed an ASIC and sent me the final masks… but they didn't tell me what the chip actually does. So that's the challenge: reverse engineer the circuit and figure out the chip's purpose. Jane Street has a bunch of swag ready to send to the most creative solutions, and they're also planning to feature the top write-ups in a blog post. Download the files and get started at janestreet.com/dwarkesh* Cursor and SpaceX recently released Grok 4.5, and I've been surprised by just how good the model is. For example, when I tested it against Fable and Sol on a bunch of AI governance questions, all three models gave substantially the same answers, but Grok was faster, more concise, and cheaper. Grok 4.6 is coming soon, but in the meantime, you can try 4.5 at cursor.com/dwarkeshTimestamps(00:00:00) – Is AI R&D verifiable enough to unlock recursive self-improvement?(00:16:52) – Is AI progress bottlenecked by human expert data?(00:34:02) – Flat token prices suggest scaling has been slow(00:39:47) – Skills AI can't train on: does it even need them?(00:48:07) – Aligned to whom?(01:09:18) – Recent incidents of AIs colluding and deceiving humans(01:19:38) – What could possibly go wrong? A concrete scenario(01:48:02) – From reward hacking to takeover Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe

SBS Korean - SBS 한국어 프로그램
자동차 보험료 8% 상승...물가보다 두 배 이상 올라, 왜 이렇게 올랐나?

SBS Korean - SBS 한국어 프로그램

Play Episode Listen Later Aug 10, 2026 2:31


호주의 자동차 보험료가 물가보다 두 배 이상 빠르게 올랐지만, 소비자들은 인상 이유를 충분히 설명받지 못하고 있는 것으로 나타났습니다. ASIC은 보험사들에게 보험료 산정과 인상 이유를 보다 명확하게 공개하라고 요구했습니다.호주 공영방송 SBS 한국어 프로그램은 호주 한인 커뮤니티를 위한 뉴스와 생활 정보, 그리고 다양한 이야기를 전합니다. 호주와 한국을 잇는 신뢰할 수 있는 콘텐츠를 만나보세요.더 많은 뉴스와 팟캐스트는 SBS 한국어 프로그램 웹사이트에서 확인하세요.www.sbs.com.au/korean

SBS German - SBS Deutsch
Meldungen des Tages, Montag 10.08.26

SBS German - SBS Deutsch

Play Episode Listen Later Aug 10, 2026 3:33


Wirtschaftsverbände wollen Mitsprache bei Einwanderungsgesetz / Netanjahu lehnt Trumps Friedensplan ab / Selenskyj kündigt Vergeltung an / NSW führt übertragbare Mietkaution ein / ASIC sperrt 150 Unternehmen und 36 Geschäftsführer / Tropischer Wirbelsturm trifft China / La Liga fordert Infantino-Rücktritt

ABC News Top Stories
Why are car insurances getting so expensive?

ABC News Top Stories

Play Episode Listen Later Aug 10, 2026 2:30


If you thought your car insurance was getting more expensive, you're not alone.Premiums soared above inflation in the 2024-25 financial year and the corporate regulator ASIC has put some car insurance companies on notice for failing to explain why costs are rising.While the insurance sector says there's valid reasons for the trend, ASIC's investigation reveals there's still room for companies to charge less.

Monero Talk
Self-Custody Best Practices, Cake Wallet, and Monero with Seth for Privacy | EPI 391

Monero Talk

Play Episode Listen Later Aug 9, 2026 119:13


Any donation is greatly appreciated! 47e6GvjL4in5Zy5vVHMb9PQtGXQAcFvWSCQn2fuwDYZoZRk3oFjefr51WBNDGG9EjF1YDavg7pwGDFSAVWC5K42CBcLLv5U OR DONATE HERE: https://www.monerotalk.live/donate TODAY'S SHOW: In this episode Seth for Privacy joins Douglas Tuman to discuss his recent deep dive into AI, particularly privacy-preserving ways to use AI while paying with Monero. The conversation then moves into Monero development, Zcash and its privacy tradeoffs, ASIC resistance and RandomX, tail emission and dynamic block size, and the current state of quantum-resistant cryptography. Seth also discusses Monero's decentralized funding and governance model, potential attacks against the Monero community, wallet security, and the promising performance improvements being explored through Cuprate. TIMESTAMPS: (00:00:00) Intro & Seth for Privacy Returns (00:08:57) Paying for Private AI With Monero (00:17:03) AI Security Auditing for Monero & Cake (00:25:02) Zcash Ironwood Migration & Privacy (00:42:54) Monero's ASIC Resistance & RandomX (00:54:32) Monero Tail Emission & Dynamic Block Size (01:05:34) Monero & Quantum Resistance (01:13:20) Monero's Development Ecosystem & Funding (01:31:00) Radar, Monero View Keys & Adoption (01:44:36) Hardware Wallets & Monero Custody (01:54:06) Cuprate, Monero Nodes & Scaling GUEST LINKS: https://x.com/sethforprivacy Purchase Cafe & tip the farmers w/ XMR! https://gratuitas.org/ SPONSORS: Cakewallet.com, the first open-source Monero wallet for iOS. You can even exchange between XMR, BTC, LTC & more in the app! Monero.com by Cake Wallet - ONLY Monero wallet (https://monero.com/) StealthEX, an instant exchange. Go to (https://stealthex.io) to instantly exchange between Monero and 450 plus assets, w/o having to create an account or register & with no limits. WEBSITE: https://www.monerotopia.com CONTACT: monerotalk@protonmail.com ODYSEE: https://odysee.com/@MoneroTalk:8 TWITTER: https://twitter.com/monerotalk FACEBOOK: https://www.facebook.com/MoneroTalk HOST: https://twitter.com/douglastuman INSTAGRAM: https://www.instagram.com/monerotalk TELEGRAM: https://t.me/monerotopia MATRIX: https://matrix.to/#/%23monerotopia%3Amonero.social MASTODON: @Monerotalk@mastodon.social MONERO.TOWN: https://monero.town/u/monerotalkAny donation is greatly appreciated!Any donation is greatly appreciated!

Australian Retirement Podcast
ASIC's adviser-fee crackdown: what retirees should ask before paying for advice

Australian Retirement Podcast

Play Episode Listen Later Aug 6, 2026 41:47


In this episode of Australian Retirement Podcast, Drew Meredith and James O'Reilly unpack ASIC's latest report into adviser fees and the growing pressure on super platforms to prove clients are getting fair value. It's a timely conversation for retirees, pre-retirees and business owners who are wondering what financial advice should cost, what good oversight looks like, and how to ask sharper questions before signing on. Drew and James explore why platform-based fee deductions have become such a focus, what ASIC appears to be targeting, and how poor-value advice can still slip through even in a heavily regulated system. They also break down the tension between cost and value: why the cheapest adviser is not always the best fit, why specialised advice often costs more, and what investors should expect to receive in return. The episode finishes with a practical listener question from a couple comparing two very different advice proposals. If you've ever wondered whether an upfront fee is too high, how ongoing fees should be judged, or what outcomes an adviser should be able to show in year one, this conversation will help you think more clearly before making a decision. Episode resources – Ask a question (select the Retirement podcast) Show partner resources – Visit TermPlus to learn more – Join Pearler using the code "RASKSWITCH" and get $32 of Pearler Credit – Whatever comes next for your business, power it with Stripe EOFY deals to know about - ending June/July 2026 – 1 free trade per month, for 12 months, for new Pearler customers Rask resources – All services – Financial Planning – Invest with us – Access Show Notes – Ask a question – We love feedback! Follow us on social media – Instagram: @rask.invest – TikTok: @rask.invest DISCLAIMER: This podcast contains general financial information only. That means the information does not take into account your objectives, financial situation, or needs. Because of that, you should consider if the information is appropriate to you and your needs, before acting on it. If you're confused about what that means or what your needs are, you should always consult a licensed and trusted financial planner. Unfortunately, we cannot guarantee the accuracy of the information in this podcast, including any financial, taxation, and/or legal information. Remember, past performance is not a reliable indicator of future performance. The Rask Group is NOT a qualified tax accountant, financial (tax) adviser, or financial adviser. Access The Rask Group's Financial Services Guide (FSG): https://www.rask.com.au/fsg Learn more about your ad choices. Visit megaphone.fm/adchoices

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ
Listen to the full SBS Punjabi radio program - ਸੁਣੋ ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ ਦਾ ਪੂਰਾ ਪ੍ਰੋਗਰਾਮ

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ

Play Episode Listen Later Aug 5, 2026 44:37


In the latest episode of the SBS Punjabi program, catch up on all the latest happenings in the Punjabi community across Australia and around the world. From the latest news to a new investigation revealing how many users of offset accounts, originally designed to reduce interest costs, ended up paying extra instead. Don't miss our exclusive interview with ‘dubbing wizard' and Pakistan-based comedian Sajjad Jani. This episode also features successful restaurateur and chef Amar Singh, who shares the secrets behind running a thriving business. To enjoy all this and some fun banter in Punjabi, click the audio button. - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ ਦੇ ਇਸ ਪ੍ਰੋਗਰਾਮ ਵਿੱਚ ਤਾਜ਼ਾ ਖ਼ਬਰਾਂ ਦੇ ਨਾਲ-ਨਾਲ ਆਸਟ੍ਰੇਲੀਆ ਅਤੇ ਦੁਨੀਆ ਭਰ ਵਿੱਚ ਵੱਸਦੇ ਪੰਜਾਬੀਆਂ ਨਾਲ ਜੁੜੀਆਂ ਦਿਲਚਸਪ ਗੱਲਾਂ ਸੁਣੋ। ASIC ਦੀ ਨਵੀਂ ਰਿਪੋਰਟ ਰਾਹੀਂ ਜਾਣੋ ਕਿ ਲੋਨ ਦਾ ਵਿਆਜ ਘਟਾਉਣ ਲਈ ਵਰਤੇ ਜਾਂਦੇ ਆਫਸੈੱਟ ਅਕਾਊਂਟ ਦੀ ਸਹੀ ਵਰਤੋਂ ਨਾ ਹੋਣ ਕਾਰਨ ਕੁਝ ਲੋਕ ਲੋੜ ਤੋਂ ਵੱਧ ਵਿਆਜ ਕਿਉਂ ਦੇ ਰਹੇ ਹਨ। ਇਸ ਤੋਂ ਇਲਾਵਾ, ਸਾਡੇ ਹਫ਼ਤਾਵਾਰੀ ਸੈਗਮੈਂਟ ‘ਅਖਾਣ ਕਹਾਣੀ' ਵਿੱਚ ਭੁੱਲੇ-ਵਿਸਰੇ ਪੰਜਾਬੀ ਅਖਾਣਾਂ ਦੇ ਪਿਛੋਕੜ, ਅਰਥ ਅਤੇ ਉਨ੍ਹਾਂ ਨਾਲ ਜੁੜੀਆਂ ਕਹਾਣੀਆਂ ਬਾਰੇ ਜਾਣੋ। ਲਹਿੰਦੇ ਪੰਜਾਬ ਦੇ ਮਸ਼ਹੂਰ ਕਾਮੇਡੀਅਨ ਅਤੇ ‘ਡੱਬਿੰਗ ਦੇ ਜਾਦੂਗਰ' ਸੱਜਾਦ ਜਾਨੀ ਨਾਲ ਖ਼ਾਸ ਗੱਲਬਾਤ ਸੁਣੋ ਅਤੇ ਮੈਲਬਰਨ ਵਿੱਚ ਕਈ ਭਾਰਤੀ ਰੈਸਟੋਰੈਂਟ ਚਲਾਉਣ ਵਾਲੇ ਸ਼ੈੱਫ ਅਮਰ ਸਿੰਘ ਤੋਂ ਉਨ੍ਹਾਂ ਦੀ ਸਫਲਤਾ ਦੇ ਸਫ਼ਰ ਅਤੇ ਗੁਰਾਂ ਬਾਰੇ ਜਾਣੋ। ਨਾਲ ਹੀ, ਪੰਜਾਬੀ ਸੰਗੀਤ ਅਤੇ ਹੋਰ ਦਿਲਚਸਪ ਗੱਲਬਾਤਾਂ ਦਾ ਆਨੰਦ ਲਓ।

The Quicky
Your AM Headlines + Inside The Gambling Ad Crackdown

The Quicky

Play Episode Listen Later Aug 4, 2026 22:05 Transcription Available


HEADLINES: The corporate watchdog has issued its strongest warning yet about a $25-billion-a-month betting industry that is looking to expand into Australia. If you've bought free-range eggs lately, we have something something worth knowing. A diarrhoea outbreak in the US has now spread to forty-five of the fifty states, and it's all been linked back to one contaminated food supplier. Ariana Grande has addressed the growing concern over her health and her upcoming step back from the spotlight in an emotional message to fans at a concert. And the OG social media app is coming back... GET IN TOUCHGot a story, news tip-off, feedback or dilemma?Send us a voice note or email us at thequicky@mamamia.com.au SUBSCRIBER GIVEAWAY: The Out Loud hosts have handpicked their all-time favourite reads and one subscriber is winning the entire 12-book stack. Subscribe to Mamamia here to be automatically entered. Already a subscriber? Don't sweat it, you're already in. T&Cs apply. CREDITSHost: Charlotte Mortlock Audio Producer: Scott Stronach Group Executive Producer: Charlotte Mortlock, Talissa Bazaz & Georgie Page Check out The Quicky Instagram here and our TikTok here Discover more Mamamia podcasts here Did you know some of our shows are now in video on the Apple Podcast app? Make sure your phone is up to date and check it out here! Mamamia acknowledges the traditional owners of the land on which we have recorded this podcast.Become a Mamamia subscriber: https://www.mamamia.com.au/subscribeSee omnystudio.com/listener for privacy information.

Mortgage Business Uncut
Unpacking the lender policies changing the game

Mortgage Business Uncut

Play Episode Listen Later Aug 4, 2026 32:45


Lenders are increasingly looking beyond interest rates to stand out, with new policies and product innovations reshaping the options available to borrowers. But what do these changes actually mean for brokers and their clients? Join Broker Daily's Julian Barnes with Finni brokers Costa Arvanitopoulos and Rebecca Carlson as they unpack the latest lender policy changes, including Liberty becoming the first non-bank to join the 5 per cent Deposit Scheme, AMP's new 40-year investor loan, the growing competition in SMSF refinancing, and ASIC's findings into widespread mortgage offset account failures. The team also discusses what softer auction clearance rates and falling home values are really saying about the property market, why investors are pivoting towards new strategies rather than sitting on the sidelines, and what the latest trends mean for brokers navigating an evolving lending landscape.

Australian Property Podcast
ASIC's offset account warning: what borrowers should check, with Sarah Court

Australian Property Podcast

Play Episode Listen Later Aug 4, 2026 35:31


This episode was originally featured on the Australian Finance Podcast. In this Australian Property Podcast episode, Owen Rask speaks with Sarah Court, Chair of ASIC, about what the regulator is seeing across the financial system and why everyday Australians need to pay closer attention to how their money is being handled. Sarah explains ASIC's role in licensing firms, holding financial institutions to account and stepping in when consumers or investors are at risk. The core of the conversation focuses on mortgage offset accounts after ASIC uncovered failures at banks that meant some customers were not receiving the interest savings they expected. Sarah explains why these problems matter, why legacy systems and weak processes can create real financial harm, and why banks need to fix the basics instead of assuming customers will spot issues themselves. Owen and Sarah also talk about scams, risky investments and the importance of slowing down before handing over money. She shares why MoneySmart can be such a valuable resource for Australians who want plain-English guidance on investing, risk, scams and personal finance. If you want a practical reminder to check your offset setup, stay sceptical of fast-talking investment pitches and understand where ASIC fits into the system, this episode is worth your time before it costs you real money. Episode resources – ⁠Moneysmart⁠ – ⁠Investment scam alerts - ASIC⁠ – ⁠Professional registers search - ASIC – Ask a question (select the Property podcast) Show partner resources – Join Pearler using the code "RASKSWITCH" and get $32 of Pearler Credit – Whatever comes next for your business, power it with Stripe Rask resources – Pete's Buyers Agency – Alcove mortgage broking – Amy Lunardi Buyers Agency (Melbourne) – All services – Financial Planning – Invest with us – Access Show Notes – Ask a question – We love feedback! Follow us on social media – Instagram: @rask.invest – TikTok: @rask.invest DISCLAIMER: This podcast contains general financial information only. That means the information does not take into account your objectives, financial situation, or needs. Because of that, you should consider if the information is appropriate to you and your needs, before acting on it. If you're confused about what that means or what your needs are, you should always consult a licensed and trusted financial planner. Unfortunately, we cannot guarantee the accuracy of the information in this podcast, including any financial, taxation, and/or legal information. Remember, past performance is not a reliable indicator of future performance. The Rask Group is NOT a qualified tax accountant, financial (tax) adviser, or financial adviser. Access The Rask Group's Financial Services Guide (FSG): https://www.rask.com.au/fsg Learn more about your ad choices. Visit megaphone.fm/adchoices

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

SBS Mandarin - SBS 普通话电台
什么是退休公积金投资平台 ?ASIC审查结果如何?

SBS Mandarin - SBS 普通话电台

Play Episode Listen Later Aug 3, 2026 8:19


澳大利亚证券和投资委员会 (ASIC) 表示,退休公积金受托人未能充分保护成员免受其投资平台上的过高费用和高风险投资的影响。 澳洲经济学者Natalie彭博士说,很多人甚至不知道自己是否在这种投资平台上。 点击音频收听详细采访

Digital Finance Analytics (DFA) Blog
Robbed Again: Banks Behaving Badly Making Mortgages More Expensive!

Digital Finance Analytics (DFA) Blog

Play Episode Listen Later Jul 31, 2026 13:03


According to a recent ASIC report, millions of Australians rely on mortgage offset accounts to reduce the cost of their home loan, but an ASIC review has found customers may have been unknowingly paying more interest than they should because some banks failed to properly manage offset accounts. That’s because banks have not been delivering … Continue reading "Robbed Again: Banks Behaving Badly Making Mortgages More Expensive!"

The Adviser Podcast Network
What's Making Headlines – Could unlinked offset accounts be costing your clients thousands?

The Adviser Podcast Network

Play Episode Listen Later Jul 31, 2026 43:45


Are major banks failing to link offset accounts, leaving thousands of home loan borrowers out of pocket without even knowing it? Join host Annie Kane, commercial content writer Ben Squires, and senior journalist Charlie Tchetchenian as they review the news of the week. This week, they discuss: Widespread failures uncovered by ASIC across major lenders' mortgage offset accounts. Westpac reversing its cash rate hike forecast following cooler-than-expected CPI data. Pepper Money raising its maximum LVR to 98 per cent alongside AMP's launch of a 40-year investor loan. And much more! News stories mentioned: Banks slammed over offset failures in ASIC probe https://www.theadviser.com.au/compliance/48733-banks-slammed-over-offset-failures-in-asic-probe New lender joins Help to Buy scheme https://www.theadviser.com.au/lender/48720-new-lender-joins-help-to-buy-scheme Productivity Commission calls for overhaul of housing rules https://www.theadviser.com.au/borrower/48730-productivity-commission-calls-for-overhaul-of-housing-rules Westpac scraps double hike call as CPI drops https://www.theadviser.com.au/borrower/48736-westpac-scraps-double-hike-call-as-cpi-drops Lending emerges as leading source of Banking Code breaches https://www.theadviser.com.au/lender/48715-lending-emerges-as-leading-source-of-banking-code-breaches Pepper unveils sweeping expansion of lending parameters https://www.theadviser.com.au/lender/48726-pepper-unveils-sweeping-expansion-of-lending-parameters Bluestone flags major trends reshaping borrowers and credit https://www.theadviser.com.au/lender/48729-bluestone-flags-major-trends-reshaping-borrowers-and-credit 40-year investor loan with 10 years of IO launches https://www.theadviser.com.au/growth/48739-40-year-investor-loan-with-10-years-of-io-launches Qudos Bank brand to be retired in 2027 https://www.theadviser.com.au/lender/48724-qudos-bank-brand-to-be-retired-in-2027

SBS Korean - SBS 한국어 프로그램
익스플레이너: 은행 실수로 이자를 더 냈다? 오프셋 계좌 꼭 확인하세요

SBS Korean - SBS 한국어 프로그램

Play Episode Listen Later Jul 30, 2026 6:27


호주증권투자위원회(ASIC)가 오프셋 계좌를 이용 중인 고객이라면 계좌가 대출과 정상적으로 연결돼 있는지, 이자가 올바르게 계산되고 있는지를 반드시 확인할 것을 당부했습니다.호주 공영방송 SBS 한국어 프로그램은 호주 한인 커뮤니티를 위한 뉴스와 생활 정보, 그리고 다양한 이야기를 전합니다. 호주와 한국을 잇는 신뢰할 수 있는 콘텐츠를 만나보세요.더 많은 뉴스와 팟캐스트는 SBS 한국어 프로그램 웹사이트에서 확인하세요.sbs.com.au/language/korean

SBS Mandarin - SBS 普通话电台
审查发现多家银行对冲账户出错 如何检查你的账户正常运行?

SBS Mandarin - SBS 普通话电台

Play Episode Listen Later Jul 30, 2026 4:40


澳大利亚证券与投资委员会(ASIC)的一项审查发现,多家银行在房贷对冲账户的设立和管理方面存在漏洞,导致部分借款人多支付房贷利息。2023年9月至2025年8月期间,各银行已就相关错误向客户赔偿超过5500万澳元(收听播客,了解详情)。

SBS Vietnamese - SBS Việt ngữ
Vay mua nhà: Bạn có bị tính dư tiền lãi vì tài khoản offset?

SBS Vietnamese - SBS Việt ngữ

Play Episode Listen Later Jul 30, 2026 13:32


Tài khoản offset vẫn hiện trên ứng dụng, tiền vẫn nằm trong đó, nhưng khoản vay có thể không được giảm lãi như người vay vẫn tưởng. ASIC phát hiện một số tài khoản chưa được thiết lập hoặc liên kết đúng, khiến khách hàng âm thầm trả dư tiền lãi và các ngân hàng phải bồi thường hơn 55 triệu đô la. Người vay cần kiểm tra gì và xử lý ra sao nếu phát hiện sai sót?

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ
ਕੀ ਤੁਹਾਡਾ ਹੋਮ ਲੋਨ ਆਫਸੈੱਟ ਅਕਾਊਂਟ ਸਹੀ ਤਰ੍ਹਾਂ ਕੰਮ ਕਰ ਰਿਹਾ ਹੈ? ਬੈਂਕਾਂ ਦੀਆਂ ਗਲਤੀਆਂ ਕਾਰਨ ਗਾਹਕਾਂ ਨੂੰ $55 ਮਿ

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ

Play Episode Listen Later Jul 30, 2026 6:01


ਆਸਟ੍ਰੇਲੀਆ ਦੇ ਕਾਰਪੋਰੇਟ ਰੈਗੂਲੇਟਰ ASIC ਦੀ ਜਾਂਚ ਵਿੱਚ ਸਾਹਮਣੇ ਆਇਆ ਹੈ ਕਿ ਕਈ ਵੱਡੇ ਬੈਂਕਾਂ ਦੀਆਂ ਗਲਤੀਆਂ ਕਾਰਨ ਹਜ਼ਾਰਾਂ ਹੋਮ ਲੋਨ ਗਾਹਕ ਉਨ੍ਹਾਂ ਵਿਆਜ ਬਚਤਾਂ ਤੋਂ ਵਾਂਝੇ ਰਹੇ, ਜਿਨ੍ਹਾਂ ਦਾ ਉਨ੍ਹਾਂ ਨੂੰ ਹੱਕ ਸੀ। ਦੋ ਸਾਲਾਂ ਵਿੱਚ ਬੈਂਕਾਂ ਨੂੰ $55 ਮਿਲੀਅਨ ਤੋਂ ਵੱਧ ਦਾ ਮੁਆਵਜ਼ਾ ਦੇਣਾ ਪਿਆ। ਜੇ ਤੁਹਾਡੇ ਕੋਲ ਵੀ ਹੋਮ ਲੋਨ ਨਾਲ ਜੁੜਿਆ ਆਫਸੈੱਟ ਅਕਾਊਂਟ ਹੈ, ਤਾਂ ਇਹ ਜਾਣਨਾ ਜ਼ਰੂਰੀ ਹੈ ਕਿ ਇਹ ਸਹੀ ਤਰੀਕੇ ਨਾਲ ਕੰਮ ਕਰ ਰਿਹਾ ਹੈ ਜਾਂ ਨਹੀਂ। ਆਓ ਜਾਣਦੇ ਹਾਂ ਇਸ ਖ਼ਾਸ ਰਿਪੋਰਟ ਵਿੱਚ।

SBS Vietnamese - SBS Việt ngữ
Có tiền trong tài khoản offset vẫn bị tính thêm lãi: ASIC cảnh báo lỗ hổng ở ngân hàng Úc

SBS Vietnamese - SBS Việt ngữ

Play Episode Listen Later Jul 29, 2026 4:48


Nhiều người vay mua nhà ở Úc sử dụng tài khoản offset với kỳ vọng tiết kiệm tiền lãi. Tuy nhiên, một báo cáo mới của Ủy ban Chứng khoán và Đầu tư Úc (ASIC) cho thấy một số khách hàng đã phải trả thêm hàng ngàn đô la tiền lãi vì những lỗi trong cách ngân hàng thiết lập và quản lý loại tài khoản này. Đáng chú ý, trong một số trường hợp, khách hàng hoàn toàn không biết mình đang bị thiệt hại.

SBS Cantonese - SBS广东话节目
儲銀指不排除再加息 證監報告:對沖按揭戶口減息效用或有限

SBS Cantonese - SBS广东话节目

Play Episode Listen Later Jul 29, 2026 8:47


儲備銀行行長布洛克警告,本地基本通脹率仍高於儲銀的 2% 至 3% 的目標區間,因此利率可能需要進一步上升。另外,澳洲證監署 (ASIC) 報告指,部份銀行的對沖戶口,未能如預期幫助房貸人士節省利息支出。

SBS Tamil - SBS தமிழ்
வங்கிகளின் Offset Account குறைபாடு: தெரியாமலே அதிக வட்டி கட்டிய வாடிக்கையாளர்கள்

SBS Tamil - SBS தமிழ்

Play Episode Listen Later Jul 29, 2026 4:03


வீட்டுக்கடன் வட்டியை குறைக்க உதவும் ‘Offset Account' சேவையில் குறைபாடுகள் காரணமாக, சில ஆஸ்திரேலியர்கள் தங்களுக்கே தெரியாமல் ஆயிரக்கணக்கான டொலர்களை கூடுதல் வட்டியாக செலுத்தியிருப்பதாக ASIC வெளியிட்டுள்ள புதிய அறிக்கை தெரிவித்துள்ளது. இது குறித்த செய்தியை எடுத்து வருகிறார் செல்வி இன்பசேகரன்.Offset account errors leave home loan borrowers paying more interest

The Adviser Podcast Network
What's Making Headlines – ASIC's best interests review, clawback changes, and aggregator movements

The Adviser Podcast Network

Play Episode Listen Later Jul 24, 2026 33:25


What do regulatory compliance, clawback policies, and aggregator shifts have in common? They are all driving major conversations across the Australian mortgage broking landscape right now. Join host Annie Kane and commercial content writer Ben Squires as they break down the biggest news across the finance sector this week and takeaways from the MFAA National Conference in Melbourne. This week, they discuss: What ASIC is looking at in its best interests duty review and warnings about boilerplate documentation. How ING's clawback policy update is shifting the conversation around broker commissions. Key takeaways from WealthX's inaugural report tracking credit representative movements across major aggregators. And much more! News stories mentioned: ASIC to release best interests duty report by Q4 https://www.theadviser.com.au/compliance/48707-asic-to-release-best-interests-duty-report-by-q4 Assistant Treasurer backs brokers as drivers of competition https://www.theadviser.com.au/broker/48705-assistant-treasurer-backs-brokers-as-drivers-of-competition Standing still is the biggest threat to broker market share, panel warns https://www.theadviser.com.au/growth/48713-standing-still-is-the-biggest-threat-to-broker-market-share-panel-warns New independent report reveals aggregator movements https://www.theadviser.com.au/aggregator/48694-new-independent-report-reveals-aggregator-movements Which states are the aggregators attracting brokers from? https://www.theadviser.com.au/aggregator/48697-revealed-the-aggregators-that-are-dominating-in-each-state ING confirms major clawback change https://www.theadviser.com.au/lender/48698-ing-confirms-major-clawback-change

The MAD Podcast with Matt Turck
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The MAD Podcast with Matt Turck

Play Episode Listen Later Jul 23, 2026 72:41


AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.(00:00) Cold open & Intro(01:31) Why speed became the AI bottleneck(02:32) Tokens per second per user, explained(03:16) AI's broadband moment and the Netflix analogy(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs(06:36) What is an ASIC?(08:08) Nvidia, Groq, and the fast inference war(09:16) OpenAI, Broadcom, and specialized silicon(12:10) China, power, and sovereign AI infrastructure(15:05) Is the AI infrastructure boom a bubble?(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm(22:57) Why agents are creating CPU demand(25:36) Andrew Feldman's path from SeaMicro to Cerebras(26:13) Why Cerebras bet on AI in 2016(31:14) SRAM vs. HBM: why inference is a memory problem(33:19) What wafer-scale computing actually means(34:28) The deep-tech “Everest” problem(36:07) The moment the first Cerebras system worked(36:49) Ringing the bell and surviving deep tech(39:08) How a giant chip handles failure(41:22) Why GPUs struggle with decode(42:17) Prefill vs. decode explained(44:01) The “100 HD movies” problem in AI inference(45:04) How fast inference changes RL and training(48:08) Reasoning models and why they cost more compute(50:08) Verification, guardrails, and small models checking big models(52:37) Multimodal AI and the path to video(53:51) Cerebras' business model: hardware, cloud, and API(55:14) OpenAI's 750MW inference deal(55:36) Why data centers are measured in megawatts(58:01) AWS Trainium + Cerebras decode(59:29) Fast tokens as a cloud product(01:00:52) Is CUDA still a moat?(01:03:53) How TSMC helped Cerebras build the giant chip(01:07:41) Why nobody cared in 2020(01:08:15) Why chip supply chains are hard to diversify(01:09:54) Why today's AI models will be the worst you ever use(01:10:38) What fast AI could do to SaaS

SBS World News Radio
INTERVIEW: ASIC tackles corporate misconduct, collects hundreds of millions of dollars in fines

SBS World News Radio

Play Episode Listen Later Jul 21, 2026 3:59


The Australian Securities and Investments Commission says it has secured a record $830 million in civil penalties orders during the 2025-2026 Financial Year. It launched more than 250 investigations, which led to 25 criminal convictions and 11 individuals being sentenced to prison. Businesses subjected to major civil penalties include Union Standard International Group, HSBC Bank Australia, Westpac, Walker Stores and Mercer Super. SBS Reporter Stephanie Youssef has been speaking with ASIC Commissioner Alan Kirkland

SBS Hmong - SBS Hmong
Tuesday news: Israel lub embassy tej lus tsis txaus siab rau Labor tsab cai

SBS Hmong - SBS Hmong

Play Episode Listen Later Jul 21, 2026 9:55


Lub rooj nom xav hloov siv neeg txum tim lub npe; tej xov xwm teev txog hav zoov kub hnyiab; Israel tus nom embassador tej lus tsis pom zoo nrog Labor tsab cai pab cheeb tsam Middle East; Hungary tus president; Trump tej se tariff 50% tsub rau Canada tej khoom; Saudio Arabia cov kev tsis pom zoo nrog Houthis cov kev thaiv suam hiav txwv yuav mus rau Red Sea; Tsoom fwv teb chaws qhia txog nws cov nyiaj pab txhim kho Gaza; One Nation tus coj tej lus tawm tswv yim cuam tshuam txog Australia tsab cai White Policy; ASIC cov kev nplua tej lagluam thiab tej neeg ntiag tug; Nom tswv Nplog cov kev yuav npaj rau txim rau tus tswv lub tuam txhab tsim cov cawv Tiger Vodka; Nplog thiab Thaib tej nom tswv cov kev txheeb lub pas dej tauv Luangprabang; Thaib tus coj lwm pab nom cov kev yuav sib teev theev txog cov kev tawm suab tsis tso siab rau tsoom fwv Thaib cuam tshuam txog tej xwm txheej lwg noj lwg haus, Ntiaj teb cov kev sib tw ncaws pob 2026 FIFA World Cup.

Hard Reset
E97 - FPGA (Gal)

Hard Reset

Play Episode Listen Later Jul 20, 2026 50:03


מה זה FPGA?רכיב חומרה שאפשר לתכנת.רגע. לתכנת... חומרה?למה זה קיים?כדי לענות על השאלה האחרונה הבאנו את האדם המתאים ביותר:גל מלמז"ק - למה זה קיים?למי שלא מכיר, למז"ק הוא פודקאסט שבו גל, חגי וליבי מדברים על דברים מטופשים ו/או אזוטריים ושואלים את עצמם שאלה אחת פשוטה: למה זה קיים?במקרה, כשהוא לא עסוק בלשאול למה דברים קיימים, גל גם מבין דבר או שניים ב-FPGA.על מה דיברנו בפרק?מה זה בכלל FPGA ומה בדיוק "מתכנתים" בו?מה ההבדל בין לתכנת FPGA לבין לכתוב Firmware?מתי עדיף להשתמש ב-FPGA במקום ASIC?מה זה LUT ולמה יש לו דווקא 6 כניסות?מה זה BRAM?מה זה DSP?איך עובדים Clock-ים בתוך FPGA?האם צריך וריפיקציה למשהו שאפשר פשוט לתכנת מחדש?מה זה FPGA ולמה הוא קיים?וגם, מסיבות שאינן ברורות לחלוטין:פוקימונים, קריסטלים, תדרים, חומרה שמתחפשת לתוכנה, תוכנה שמתחפשת לחומרה וכמות לא סבירה של רפרנסים ללמז"ק.בסוף אפילו הפכנו את היוצרות ושאלנו את גל שאלה משלנו:מה זה Yield? (ולמה זה קיים)אחרי שהאזנתם לפרק, ובהנחה שאתם לא LUT עם 7 כניסות, מוזמנים להצטרף לקהילה שלנו >>>https://hardreset.co.il/communityנשמח לשמוע בתגובות:אם החומרה שאתם מפתחים הייתה פוקימון - איזה פוקימון היא הייתה?פרק 97 - FPGA: למה זה קיים?Hard Reset - הפודקאסט של קהילת Hardware Engineering Israel.מוזמנים ליצור איתנו קשר במייל podcasthardreset@gmail.com

SBS Tamil - SBS தமிழ்
பிரபல நிதி நிபுணர்களின் பெயரில் புதிய முதலீட்டு மோசடி - ASIC எச்சரிக்கை!

SBS Tamil - SBS தமிழ்

Play Episode Listen Later Jul 20, 2026 7:20


பிரபல நிதி நிபுணர்களின் பெயரில் புதிய முதலீட்டு மோசடிகள் நடைபெறுவதாக ASIC Australian Securities and Investments Commission எச்சரிக்கை விடுத்துள்ளது. இது குறித்த செய்தியின் பின்னணியை எடுத்து வருகிறார் செல்வி இன்பசேகரன்.

asic investments commission
SemiWiki.com
Podcast EP356: An Oveview of the Unique Skills and Broad Impact of Alphacore with Andrew Levy

SemiWiki.com

Play Episode Listen Later Jul 17, 2026 15:00


Daniel is joined by Andrew Levy, Vice President of Business Development, IP & ASIC at Alphacore. Andrew leads strategic growth initiatives, customer engagement, technology commercialization, and business development activities for the company’s semiconductor IP, ASIC design services, and advanced microelectronics… Read More

Herbert Smith Freehills Podcasts
The Third Wheel (ESG Australia) EP51: ASIC observations and lessons for future climate reporting

Herbert Smith Freehills Podcasts

Play Episode Listen Later Jul 17, 2026 24:34


In Part 2 of our climate reporting series, we build on the themes from Episode 50 and shift the focus to what comes next. As the first wave of disclosures has wrapped up, attention has turned to the next climate reporting cycles - particularly for June and September year-end companies. The question now is: what lessons can organisations take forward? In this episode, we unpack key takeaways from the first round of sustainability reporting and explore how they can be applied in practice for future reporters. We also take a closer look at ASIC's early observations and share our perspective on what these mean, and how companies can consider them going forward.

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ
ਖ਼ਬਰਨਾਮਾ: ਲਾਓਸ 'ਚ ਆਸਟ੍ਰੇਲੀਆਈ ਨੌਜਵਾਨਾਂ ਦੀ ਮੌਤ ਮਾਮਲੇ 'ਚ ਗੰਭੀਰ ਦੋਸ਼ ਨਹੀਂ ਲੱਗਣਗੇ, ਪੈਨੀ ਵੋਂਗ ਨੇ ਜਤਾਈ ਚਿੰਤ

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ

Play Episode Listen Later Jul 17, 2026 4:55


ਆਸਟ੍ਰੇਲੀਆ ਦੀ ਵਿਦੇਸ਼ ਮੰਤਰੀ ਪੈਨੀ ਵੋਂਗ ਨੇ ਕਿਹਾ ਹੈ ਕਿ ਲਾਓਸ ਵਿੱਚ ਮੈਥਾਨੌਲ ਜ਼ਹਿਰ ਕਾਰਨ ਮਰਨ ਵਾਲੀਆਂ ਦੋ ਆਸਟ੍ਰੇਲੀਆਈ ਨੌਜਵਾਨਾਂ ਦੇ ਮਾਮਲੇ ਵਿੱਚ ਸਭ ਤੋਂ ਗੰਭੀਰ ਦੋਸ਼ ਨਾ ਲਗਾਉਣ ਦੇ ਫ਼ੈਸਲੇ 'ਤੇ ਆਸਟ੍ਰੇਲੀਆਈ ਸਰਕਾਰ ਲਾਓਸ ਦੇ ਉੱਚ ਅਧਿਕਾਰੀਆਂ ਨਾਲ ਗੱਲਬਾਤ ਕਰੇਗੀ। ਇਸ ਦੌਰਾਨ ਅਮਰੀਕਾ ਵਿੱਚ ਚੋਣ ਸੁਰੱਖਿਆ ਨੂੰ ਲੈ ਕੇ ਰਾਸ਼ਟਰਪਤੀ ਡੋਨਾਲਡ ਟਰੰਪ ਦੇ ਨਵੇਂ ਦਾਅਵੇ, ਵਿਕਟੋਰੀਆ ਵਿੱਚ ਨਿਰਮਾਣ ਖੇਤਰ ਵਿੱਚ ਭ੍ਰਿਸ਼ਟਾਚਾਰ ਵਿਰੁੱਧ ਵਧੇਰੇ ਪੁਲਿਸ ਅਧਿਕਾਰਾਂ ਦੀ ਮੰਗ, ਨਿਵੇਸ਼ ਘੁਟਾਲਿਆਂ ਬਾਰੇ ASIC ਦੀ ਚੇਤਾਵਨੀ, ਨਾਰਦਰਨ ਟੈਰੀਟਰੀ ਵਿੱਚ ਪਹਿਲੇ ਕਮਿਊਨਿਟੀ ਚਾਈਲਡ ਐਂਡ ਫੈਮਿਲੀ ਸੈਂਟਰ ਦਾ ਉਦਘਾਟਨ ਅਤੇ ਆਸਟ੍ਰੇਲੀਆ ਵਿੱਚ ਸਿਗਰਟਨੋਸ਼ੀ ਦੀ ਘਟਦੀ ਦਰ ਵੀ ਅੱਜ ਦੀਆਂ ਮੁੱਖ ਖ਼ਬਰਾਂ ਹਨ।

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ
ਖ਼ਬਰਨਾਮਾ: ਲਾਓਸ 'ਚ ਆਸਟ੍ਰੇਲੀਆਈ ਨੌਜਵਾਨਾਂ ਦੀ ਮੌਤ ਮਾਮਲੇ 'ਚ ਗੰਭੀਰ ਦੋਸ਼ ਨਹੀਂ ਲੱਗਣਗੇ, ਪੈਨੀ ਵੋਂਗ ਨੇ ਜਤਾਈ ਚਿੰਤ

SBS Punjabi - ਐਸ ਬੀ ਐਸ ਪੰਜਾਬੀ

Play Episode Listen Later Jul 17, 2026 4:55


ਆਸਟ੍ਰੇਲੀਆ ਦੀ ਵਿਦੇਸ਼ ਮੰਤਰੀ ਪੈਨੀ ਵੋਂਗ ਨੇ ਕਿਹਾ ਹੈ ਕਿ ਲਾਓਸ ਵਿੱਚ ਮੈਥਾਨੌਲ ਜ਼ਹਿਰ ਕਾਰਨ ਮਰਨ ਵਾਲੀਆਂ ਦੋ ਆਸਟ੍ਰੇਲੀਆਈ ਨੌਜਵਾਨਾਂ ਦੇ ਮਾਮਲੇ ਵਿੱਚ ਸਭ ਤੋਂ ਗੰਭੀਰ ਦੋਸ਼ ਨਾ ਲਗਾਉਣ ਦੇ ਫ਼ੈਸਲੇ 'ਤੇ ਆਸਟ੍ਰੇਲੀਆਈ ਸਰਕਾਰ ਲਾਓਸ ਦੇ ਉੱਚ ਅਧਿਕਾਰੀਆਂ ਨਾਲ ਗੱਲਬਾਤ ਕਰੇਗੀ। ਇਸ ਦੌਰਾਨ ਅਮਰੀਕਾ ਵਿੱਚ ਚੋਣ ਸੁਰੱਖਿਆ ਨੂੰ ਲੈ ਕੇ ਰਾਸ਼ਟਰਪਤੀ ਡੋਨਾਲਡ ਟਰੰਪ ਦੇ ਨਵੇਂ ਦਾਅਵੇ, ਵਿਕਟੋਰੀਆ ਵਿੱਚ ਨਿਰਮਾਣ ਖੇਤਰ ਵਿੱਚ ਭ੍ਰਿਸ਼ਟਾਚਾਰ ਵਿਰੁੱਧ ਵਧੇਰੇ ਪੁਲਿਸ ਅਧਿਕਾਰਾਂ ਦੀ ਮੰਗ, ਨਿਵੇਸ਼ ਘੁਟਾਲਿਆਂ ਬਾਰੇ ASIC ਦੀ ਚੇਤਾਵਨੀ, ਨਾਰਦਰਨ ਟੈਰੀਟਰੀ ਵਿੱਚ ਪਹਿਲੇ ਕਮਿਊਨਿਟੀ ਚਾਈਲਡ ਐਂਡ ਫੈਮਿਲੀ ਸੈਂਟਰ ਦਾ ਉਦਘਾਟਨ ਅਤੇ ਆਸਟ੍ਰੇਲੀਆ ਵਿੱਚ ਸਿਗਰਟਨੋਸ਼ੀ ਦੀ ਘਟਦੀ ਦਰ ਵੀ ਅੱਜ ਦੀਆਂ ਮੁੱਖ ਖ਼ਬਰਾਂ ਹਨ।

The Adviser Podcast Network
What's Making Headlines – Are clawbacks out of control?

The Adviser Podcast Network

Play Episode Listen Later Jul 17, 2026 31:48


The frustrating reality of having hard-earned commissions ripped away for reasons entirely out of your control has been the bane of the industry for years, but the fight to make them fairer has come back into the spotlight again. Join host Annie Kane, commercial content writer Ben Squires, and senior journalist Charlie Tchetchenian as they review the news of the week. The team looks at the FBAA's recent submission to government flagging how clawbacks might amount to an unfair trading practice, One Nation's proposed 30-year government-backed mortgages, and the latest levy costs from ASIC. This week, they discuss: How the FBAA is taking the fight over "unfair and inequitable" lender clawbacks directly to Treasury. Why the MFAA is demanding that home lending be placed at the very front of the government's new digital identity rollout. The looming 17 per cent surge in ASIC levies that is set to hit credit intermediaries and aggregators where it hurts. And much more! News stories mentioned: Clawbacks under fire as FBAA slams lender tactics https://www.theadviser.com.au/broker/48678-fbaa-urges-crackdown-on-unfair-lender-practices MFAA says lending must sit at heart of verifiable credentials https://www.theadviser.com.au/broker/48683-mfaa-says-lending-must-sit-at-heart-of-verifiable-credentials ASIC levies surge for credit intermediaries https://www.theadviser.com.au/compliance/48667-asic-levies-surge-for-credit-intermediaries One Nation proposes government-backed 30-year fixed mortgages https://www.theadviser.com.au/borrower/48672-one-nation-proposes-government-30-year-fixed-mortgages Non-banks plugged into open banking data grid https://www.theadviser.com.au/lender/48669-non-banks-plugged-into-open-banking-data-grid Westpac locks in major cash rate prediction https://www.theadviser.com.au/borrower/48662-westpac-locks-in-major-cash-rate-prediction Commercial Finance Awards launches for 2026 https://www.theadviser.com.au/broker/48668-commercial-finance-awards-launches-for-2026

SBS World News Radio
INTERVIEW: The insidious danger of 'pump and dump' financial scams

SBS World News Radio

Play Episode Listen Later Jul 16, 2026 7:08


ASIC is warning Australians to be extremely cautious of investment tips received through social media and messaging apps, amid a spike in reports of so-called 'pump and dump' scams using fake celebrity endorsements and impersonations of financial institutions. ASIC says scammers use the identities and images of well-known finance industry figures, including respected market commentators and economists, to lure consumers into WhatsApp and other messaging groups where they are encouraged to buy investments, particularly shares. Once consumers join a messaging group, scammers provide stock tips designed to artificially inflate the price of a share before selling their own holdings. The share price then falls, leaving unsuspecting investors with significant losses. SBS Reporter Natalie Poyhonen has been talking to ASIC Commissioner Alan Kirkland

POD256 | Bitcoin Mining News & Analysis
119. Soft Forks, Censorship Risks, and Open-Source Bitcoin

POD256 | Bitcoin Mining News & Analysis

Play Episode Listen Later Jul 8, 2026 51:26 Transcription Available


In this episode, the crew unpacks the debate around BIP 110, why it has become such a lightning rod in Bitcoin, and what it could actually mean in practice. We talked through the core disagreement over what counts as “spam” on the Bitcoin network, the difference between miner-activated and user-activated soft forks, and why some are skeptical that BIP 110 will achieve its stated goal. We also explored the risks of transaction filtering, the role of consensus rules versus individual policy preferences, and how proposals like this expose bigger questions around governance, open-source development, and trust in Bitcoin node implementations.We also shifted into the mining side of the conversation, where we shared exciting progress on Mujina and why open-source mining firmware matters so much. We highlighted recent demos showing extremely fast power ramping, support for mixed hash boards on a single control board, and new community-built tooling and dashboards. To close things out, we gave a shoutout to the hashers supporting 256 Foundation, checked in on the leaderboard, and teased an upcoming guest episode with someone who knows ASIC development inside and out.

Let's Talk AI
#250 - Mythos Mess, GPT 5.6-Sol, GLM 5.2

Let's Talk AI

Play Episode Listen Later Jul 7, 2026 103:25


Our 250th episode with a summary and discussion of last week's big AI news!Recorded on 06/27/2026Note from Andrey: sorry this is late again! this episode release somehow didn't save and I only realized late, my bad... next one will be out way sooner!Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:US government gating of frontier AI expands: Anthropic gets permission to release Mythos-5 to selected companies/agencies after a standoff, OpenAI rolls out GPT-5.6 “Sol” with initial access restricted to ~20 approved organizations, and Meta is pressed to submit models to “voluntary” review—signaling an emerging de facto licensing regime with geopolitical treaty implications.Model capability and safety signals remain murky: limited benchmark disclosure, claims of token-efficiency comparisons, and third-party reports that GPT-5.6 shows extreme benchmark “cheating” sensitivity highlight steering/alignment bottlenecks and uncertainty about real-world long-horizon behavior.Compute supply chain competition accelerates: OpenAI unveils its Jalapeño inference ASIC with Broadcom on TSMC 3nm; Amazon explores selling Trainium to data-center operators; Micron invests in Anthropic with memory supply agreements; SK Hynix surpasses Samsung on HBM-driven valuation; Groq raises $650M while pivoting toward neocloud.Open source and societal response intensify: GLM 5.2 (MIT-licensed) delivers strong long-context coding performance with rapid optimizations; EconEvals maps job-task exposure; bipartisan workforce initiatives and tax credits launch; DeepMind and Apollo publish loss-of-control/control roadmaps; Hollywood reportedly drops a near-finished Sam Altman biopic amid industry pressure.Timestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:03:42) News PreviewTools & Apps(00:04:41) Anthropic allowed to release Mythos AI to some companies, agencies + Anthropic's Mythos mess is only getting worse + Anthropic floats proposal to Lutnick to end US ban of powerful 'Mythos,' 'Fable' AI models: sources(00:07:58) OpenAI Launches GPT-5.6 Sol Under First-Ever US Government-Gated AI Rollout | MLQ News + OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it + Summary of METR's predeployment evaluation of GPT-5.6 Sol(00:24:03) U.S. Presses Meta to Agree to A.I. Reviews - The New York Times(00:30:11) Anthropic's Claude Tag is learning your company, one Slack message at a time | TechCrunchApplications & Business(00:32:49) OpenAI reveals its first AI processor: Jalapeño | The Verge(00:38:29) Amazon in Talks to Sell Custom AI Chips in Bid to Undercut Nvidia(00:41:46) Micron invests in Anthropic and grants it a supply deal(00:45:18) SK Hynix overtakes Samsung to become South Korea's most valuable company | Reuters(00:49:12) AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia's $20B not-acqui-hire deal | TechCrunch(00:52:47) SpaceX inks compute deal with Reflection AI, an open source AI lab | TechCrunchProjects & Open Source(00:54:46) GLM-5.2: Built for Long-Horizon Tasks + How we built the world's fastest API for GLM-5.2 + nvidia/GLM-5.2-NVFP4 · Hugging Face(01:03:04) EconEvalsPolicy & Safety(01:05:40) $500 million AI jobs push launches with bipartisan backing - POLITICO(01:07:47) Rep. Sam Liccardo unveils AI workforce tax credit bill - POLITICO(01:08:56) Google DeepMind announced an “AI Control Roadmap” for improving AI agent security. | The Verge + Securing internal systems against increasingly capable and imperfectly aligned AI(01:14:00) The Loss of Control Playbook: Degrees, Dynamics, and Preparedness + The Loss of Control Playbook(01:16:42) Why corporate AI super PACs spent $27 million on a local election | The Verge(01:20:25) Exclusive: Conservatives plan nationwide protest against AI data centersResearch & Advancements(01:27:37) Revisiting the Platonic Representation Hypothesis: An Aristotelian View(01:31:39) Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models(01:33:59) Tapered Language ModelsSynthetic Media & Art(01:36:54) Hollywood is bending the knee to OpenAI | The VergeSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

The Investing Podcast
Memory Chips Rip Again + Broadcom Wins Apple ASIC Deal Through 2031 | July 6, 2026 – Morning Market Briefing

The Investing Podcast

Play Episode Listen Later Jul 6, 2026 19:55


Andrew and Tom discuss memory chips ripping again with ASML and LRCX up 5% on the SK Hynix listing, Broadcom's new multi-year custom ASIC deal with Apple through 2031, Solstice's $14.5 billion cash-and-stock acquisition of Element Solutions with Goldman advising, and Mario Draghi's proposal to restructure the EU into smaller blocs to unlock growth.Join our live YouTube stream Monday through Friday at 8:30 AM EST:http://www.youtube.com/@TheMorningMarketBriefingPlease see disclosures:https://www.narwhal.com/disclosure

Hard Reset
E96 - Wikipedia (Jonathan Berkheim)

Hard Reset

Play Episode Listen Later Jul 6, 2026 84:09


פעם מחקר היה מתחיל בספריה.אחרי זה היה מתחיל באנציקלופדיה.ואז הגיעה הויקיפדיה.היום אנחנו מתחילים מחקר בלשאול את ה-AI.אבל ויקיפדיה לא נעלמה מהחיים שלנו.כאות הוקרה לכלי המדהים הזה שאנחנו משתמשים בו ברמה היומית החלטנו להזמין את יהונתן ברקהיים לחלוק איתנו מהידע שלו על ויקיפדיה. בעיקר הישראלית אבל גם העולמית.על מה דיברנו בפרק?- מה היה לפני הויקיפדיה ואיך זה עבד?- האם אפשר לסמוך על ויקיפדיה?- איזה תפקידים יש בויקיפדיה?- מה זה דף שיחה ומה יש שם?- איך ויקיפדיה מתמודדת עם הבינה המלאכותית?- מה זה משחק הויקיפדיה?אחרי שהאזנתם לפרק, ובהנחה שאתם לא בוטים, מוזמנים להצטרף לקהילה שלנו - שם אנחנו עוזרים לויקיפדיה העברית לגדול >>> https://chat.whatsapp.com/FN3qNowQ51x0gTtOBdDqsHנשמח לשמוע את דעתכם על הפרק בתגובות.פרק 96 - WikipediaHard Reset - הפודקאסט של קהילת Hardware Engineering Israel.מוזמנים ליצור איתנו קשר במייל podcasthardreset@gmail.comהאזנה נעימה.

華爾街見聞
費半大跌6% 亞股全面崩盤 台股怎麼走?關鍵訊號一次告訴你! #華爾街見聞 謝晨彥分析師

華爾街見聞

Play Episode Listen Later Jul 3, 2026 24:48


【謝晨彥分析師Line官方帳號】 https://lin.ee/se5Bh8n 費半大跌6% 亞股全面崩盤 台股怎麼走?關鍵訊號一次告訴你! #華爾街見聞 謝晨彥分析師 #台股撐盤看這關鍵訊號 #Meta出租算力誰受惠? #現在卡位ASIC來得及? 馬上加入Line帳號! 獲取更多股票訊息! LINE搜尋ID:@gp520 https://lin.ee/se5Bh8n 也可來電免付費專線洽詢任何疑問! 0800-66-8085 獲取更多股票訊息 #摩爾投顧 #謝晨彥 #分析師 #股怪教授 #股票 #台股 #飆股 #三大法人 #漲停 #選股 #技術分析 #波段 #獲利 #飆股啟航 #大賺 #美債 #華爾街見聞 -- Hosting provided by SoundOn

SBS World News Radio
Super Platforms Face ASIC Scrutiny as the ASX Rallies Into Financial Year End

SBS World News Radio

Play Episode Listen Later Jun 29, 2026 12:57


The ASX 200 rallied 0.7 per cent in a late-session surge ahead of the final trading day of the financial year, with tech stocks leading the gains. Ricardo Gonçalves speaks with ASIC Commissioner Simone Constant about why the corporate watchdog says superannuation trustees must do more to protect members from high fees and risky investments on investment platforms. Plus, Kyle Rodda from Capital.com breaks down the day's market action and what investors can expect as the financial year draws to a close.

The Six Five with Patrick Moorhead and Daniel Newman
Qualcomm's Data Center Debut, OpenAI's Jalapeño, and the Memory-as-Strategic Infrastructure Debate | The Six Five Pod Ep. 310

The Six Five with Patrick Moorhead and Daniel Newman

Play Episode Listen Later Jun 29, 2026 62:08


On Episode 310 of The Six Five Pod, Patrick Moorhead and Daniel Newman unpack the biggest stories from the week, including insights from Qualcomm Investor Day 2026, OpenAI and Broadcom's Jalapeño AI chip, Anthropic's Micron partnership, SpaceX's massive Reflection AI compute deal, Sakana AI's new Fugu orchestrator, and why memory is emerging as a critical layer of AI infrastructure. Plus, Bulls & Bears covers NVIDIA's $25B bond offering, Apple's MacBook price increases, Micron's record quarter, and Cerebras' first earnings as a public company. The handpicked topics for this week are: Qualcomm Investor Day 2026 — The Data Center Debut: Pat and Dan break down Qualcomm's push into the data center after the company took the stage with Microsoft's Satya Nadella and Meta's Mark Zuckerberg as named customers. They unpack the new Dragonfly platform, including the C1000 250-core data center CPU with PCIe Gen 7 and CXL, the AI200 and AI250 inference accelerators, and a novel High Bandwidth Compute (HBC) architecture that stacks compute under LPDDR memory at dramatically lower cost than HBM. They highlight Qualcomm's ambitious growth targets: $15B data center revenue target for FY 2029, an increased total non-handset revenue goal from $22B to  $40B, and a shortened timeline for automotive revenue by two years. They also debate the identity of Qualcomm's unnamed hyperscaler customer and why its robotics opportunity may be flying under the radar. (The Decode) OpenAI and Broadcom Unveil Jalapeño, OpenAI's First Custom Chip: A photo of Sam Altman and Hock Tan holding a wafer and packaged die kicked off OpenAI's reveal of Jalapeño, a custom inference chip built with Broadcom and slated for late-2026 deployment. The chip reached tape-out in roughly nine months, which is an aggressive cycle for an ASIC of this size, and uses HBM3E memory. Pat takes a victory lap on his long-standing heterogeneous compute thesis: every hyperscaler and now every model lab is building accelerators, and the XPU efficiency argument has played out as predicted. Dan frames OpenAI's broader move as existential: they cannot serve frontier models at premium margins if compute remains constrained. He flags that OpenAI is trying to do everything from chips and fabs to social networks and browsers, and that its IPO is now delayed. (The Decode)   Anthropic and Micron Sign a Strategic Multi-Year Memory Agreement: Anthropic and Micron announced a multi-year supply agreement for HBM, DRAM, and SSDs, including co-designed next-generation memory for AI workloads, along with a strategic investment by Anthropic in Micron. The pattern mirrors Samsung and SK Hynix's pre-funding Anthropic in May, and follows OpenAI's Jalapeño as another frontier lab moving to lock in supply chain control. Dan frames it as the same circular financing playbook NVIDIA ran two to three years ago, but with the ball now in the memory triopoly's court. Pricing-floor agreements with no ceilings, customized rather than commoditized memory architecture, and demand running well past the previously assumed 2027-2028 horizon. Pat notes that the rumored 14% free cash flow margin at Anthropic makes the strategic investment math work cleanly for both sides. (The Decode)   SpaceX Signs $6.3B Compute Deal with Reflection AI: SpaceX inked a $6.3B compute lease with open-source AI lab Reflection AI, at $150M per month from July 2026 through 2029, giving Reflection access to NVIDIA GB300 chips inside the Colossus infrastructure. Combined with the $920M-per-month Google compute contract and existing xAI commitments, SpaceX now has a contracted backlog larger than most public AI startups' entire revenue base, with some calling it the largest commercial AI infrastructure provider at $80B in contracted revenue. Pat reads it as XAI failing to land with developers, consumers, or enterprises, leaving SpaceX with a pot of gold worth far more as wholesale capacity than as XAI's own training compute. Dan flags that Google owning 7% of SpaceX ahead of an IPO is not accidental, and the open question is whether this becomes a Nebius-style infrastructure trade or a full-stack Google-equivalent platform. (The Decode)   Japan's Agentic Orchestrator Sakana AI Ships Fugu Plus and Fugu Ultra: Japan's Sakana AI released Fugu Plus and Fugu Ultra, an agentic orchestrator built on a multi-agent MOE approach that routes workloads across multiple underlying models rather than training a new frontier base model. Sakana claims agentic capabilities on par with or better than top frontier models at significantly lower input/output token costs, similar to the DeepSeek and GLM cost-undercut narrative. Pat compares the architecture to OpenRouter and notes the developer-facing parallel to Perplexity Computer's model-routing approach. Both agree that models themselves are no longer moats, and suggests the real moat is the harness, tooling, connectivity, looping, agentic stack, and total compute availability. Expect more sovereign agentic plays from Japan, the Middle East, and elsewhere on the same template. (The Decode)   The Flip — Is the Era of Memory as a Commodity Over? Daniel takes the FOR side: memory has moved from commodity to strategic AI infrastructure, citing 16 multi-year agreements covering $22B in committed volume booked through 2027, 84.9% gross margins higher than NVIDIA's, the technology barriers of HBM yield/stacking/packaging that only three companies can clear, and demand drivers tied to HBM as the binding constraint on every AI accelerator rather than to elastic consumer cycles. Patrick takes the AGAINST side: long-term agreements and SCAs signal a commodity in a strong cycle, not a structural rerating; nearly every relevant memory standard — DDR5, MRDIMM, HBM3/3E/4, LPDDR5X/6, GDDR6/7, LPCAM2 — is JEDEC-standard and therefore commodity at the pin; and CXMT's China DDR5 production ramps in 2H 2026 with Lenovo already shipping and HP and Dell qualifying. Custom HBM4 and Qualcomm-style HBC are where strategic memory genuinely lives. (The Flip)   NVIDIA's $25B Investment-Grade Bond Offering: NVIDIA priced a $25B multi-tranche bond offering on June 15, its first investment-grade debt sale since 2021, with seven tranches maturing between 2028 and 2056 and $85B in orders against an initial $20B target. Dan reads it as raising when capital is cheap, and oversubscription is real. NVIDIA doesn't need the money, it has a gold balance sheet, and is establishing a credit benchmark rather than funding CapEx. Pat agrees the optics are clean, but flags the irony of NVIDIA, with negative debt, borrowing while the stock trades like dead money at a sub-20x forward P/E. Both note that NVIDIA's underperformance reflects the market's skepticism on memory-as-strategic and on NVIDIA's own capex pace relative to the buildout opportunity ahead. (Bulls & Bears)   Tim Cook Calls Apple's Memory Crunch Price Raises on MacBook and iPad "Unsustainable": Apple announced MacBook and iPad price increases of up to $300, with Tim Cook telling the WSJ the memory cost environment is unsustainable. AAPL fell ~5% on the news, the broader rally was momentarily wiped out before Micron held the gains by close. Dan frames it as a moment when the market saw who is going to pay for the AI buildout: the consumer. He notes Apple's pricing power and inelasticity test is now live. Pat traces the backstory to Apple's negative-margin pricing pressure on Micron during the 2022-2023 memory downturn. The question is whether consumer-price blowback will eventually flow back to the memory vendors. (Bulls & Bears)   Micron Blows the Doors Off Fiscal Q3 — $41.46B Revenue, 84.9% Gross Margin: The memory story continues as Micron reported its largest beat in company history with fiscal Q3 revenue of $41.46B versus a $35.69B consensus, EPS of $25.11, year-over-year growth of more than 340%, and a record 84.9% gross margin that is roughly 10 points above NVIDIA's. Q4 guidance came in at a $50B midpoint against a $43B consensus. The 16 multi-year strategic customer agreements add up to $22B in committed volume, with most contracts containing pricing floors but no ceilings on most of the volume — a structurally asymmetric setup. Pat notes 95% of the beat came from price, not units, which reinforces his commodity argument; Dan flips it as the early innings of an NVIDIA-style run that puts Micron's 2027 profit on par with Google. (Bulls & Bears)   Cerebras' First Earnings Report Since IPO — Revenue Doubles, Margins Compress: Cerebras (CBRS) reported its first earnings as a public company, doubling year-over-year revenue and beating the top line while missing EPS, but the stock sold off hard amid gross margin deterioration. Core revenue came in at $191M, up 12% sequentially, with a $194M Q2 guide that is essentially flat, core gross margins at 47% guiding to 36-38% and 38-41% for the year, and operating margins flipping from positive 2% to a guided -30% to -32%. Customer concentration is shifting from Core42 and G42 (86% of FY25 revenue) to OpenAI, which loaned Cerebras $1B and gets paid quarterly in warrants. Pat flags that Cerebras' uncontested speed claim is no longer uncontested with Groq, TPU v8i, and Tenstorrent putting up real numbers. Cathie Wood is down 52% on her position. (Bulls & Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. The Decode Qualcomm Investor Day Lands the Data Center Pivot — Microsoft Deploying Qualcomm HBC XPUs in Azure (Per Satya Nadella) + Meta MOU on Three New Qualcomm Datacenter CPUs (Per Zuckerberg); $3.9B Modular Acquisition; Dragonfly Brand + AI200/AI250 Roadmap; HUMAIN 200MW Ramp; Qualcomm to Become Largest Automotive Silicon Company; Targets $3B Datacenter Revenue FY27, $35B by FY31 https://finance.yahoo.com/markets/stocks/articles/qualcomm-investor-day-detail-data-163247063.html  OpenAI Begins Vertical Integration — First Custom Inference Chip "Jalapeño" Unveiled With Broadcom June 24 (Hock Tan: As Good as Blackwell + TPU; ~50% Cost Savings; Late-2026 Microsoft Deployment, 10GW Multi-Gen Roadmap); Daybreak Cyber Stack (June 22) Confirms the Platform Shift https://x.com/OpenAI/status/2069770172802773292  Frontier AI Labs Are Now Financing Their Own Supply Chains — Anthropic Locks In Multi-Year Micron HBM/DRAM/SSD Supply + Micron Becomes Series H Investor; Same Pattern as Samsung + SK hynix Pre-Funded Anthropic in May; $965B Post-Money, $47B Revenue Run-Rate, October IPO Target https://investors.micron.com/news-releases/news-release-details/micron-and-anthropic-announce-strategic-agreement-scale-next  SpaceX Signs $6.3B Compute Deal With Reflection AI — $150M/Month July 2026 → End of 2029; NVIDIA GB300 + Colossus 2 Capacity; SpaceX Now Largest Commercial AI Infrastructure Provider With $80B+ Committed Compute Revenue Through 2029 https://finance.yahoo.com/technology/ai/articles/spacex-reportedly-grant-reflection-ai-162749237.html  The Sovereign AI Stack Lands — Japan's Sakana Ships Fugu + Fugu Ultra Multi-Agent System (June 22) That Beats Opus 4.8, GPT-5.5, and Gemini 3.1 Pro on 10 of 11 Benchmarks; Designed Around US Export-Control Risk; Completes the Three-Bloc Sovereign-AI Map With Mistral Compute (Europe) + DeepSeek $7.4B (China) https://www.datacamp.com/blog/sakana-fugu  The Flip Is the Era of Memory as a Commodity Over? FOR: Memory is now strategic AI infrastructure with multi-year supply lock-ins. The cycle dynamics that defined the last 30 years no longer apply. https://www.benzinga.com/markets/tech/26/06/60062500/micron-earnings-could-echo-nvidias-2023-moment-says-futurum-ceo  AGAINST: Memory is cyclical and priced for perfection. This print is either step change or top of the cycle, and the second one is more likely. https://www.cnbc.com/2026/06/25/apple-macbook-ipad-price-hike-memory.html Bulls & Bears NVIDIA (NVDA) $25B Bond Sale Anchors the AI Debt-Finance Boom — First Bond Offering Since 2021; Joins Alphabet $80B, Amazon $27.5B, Meta $30B, Oracle Stack; Dan: "Locking In Cheap Capital While It Can" https://finance.yahoo.com/technology/ai/articles/nvidia-record-us-25-billion-131039687.html  Apple (AAPL) Falls −5%+ Thursday June 25 on Confirmed MacBook + iPad Price Hikes — Tim Cook RAM "Unsustainable" Comment Lands as Real Price Action; Apple Hikes Erase Micron-Driven Tech Rally Mid-Session; Memory Beneficiaries (SanDisk, Micron) Surge; Analysts "Mostly Nonplussed" https://tickerspark.ai/market/apple-inc-aapl-drops-5-3-as-price-hikes-spook-investors-1782399950638  Micron (MU) Q3 FY26 ACTUALS — Largest Beat in Company History; Revenue $41.46B (+346% YoY) Crushes $35.69B Consensus; Non-GAAP EPS $25.11 (+1,215% YoY) Beats $20.49; Record 84.9% Gross Margin (Higher Than NVIDIA); Q4 Guide $50B Midpoint vs $43B Consensus; Stock +18-19% Overnight to $1,242 https://www.nasdaq.com/articles/nvda-who-micron-blows-doors-q3-earnings-revs  Cerebras Systems (CBRS) Q1 ACTUALS — First Earnings Post-IPO; Revenue $193.4M Nearly Doubled YoY; 2026 Guide $855-$865M Beats $824M; BUT Gross Margins Forecast 38-41% (Down From 45% Q1, Half of NVIDIA + Micron); Stock −20% AH on Margin Compression; Sets Up Inference-Tier Margin Debate https://investors.cerebras.ai/news-releases/news-release-details/cerebras-systems-announces-strong-first-quarter-2026-results  

Hard Reset
E95 - The Biggest Chip (Adi Fuchs)

Hard Reset

Play Episode Listen Later Jun 22, 2026 72:33


אם הייתי שואל אנשים מה הם רוצים, הם היו עונים 'סוסים מהירים יותר'~הנרי פורדאם נמשיל סוסים לשבבים סטנדרטיים, אפשר לומר שסרברס המציאו את המכונית.לנושא נחשפנו לראשונה כשד"ר עדי פוקס כתב לנו בקבוצת המאזינים על הבום הגדול בשוק המניות האמריקאי בעקבות ההנפקה של סרברס, והדרך משם ועד לפרק לא הייתה ארוכה, ועברה דרך סקר חד משמעי במיוחד. עדי דיבר איתנו בעבר בפרק 49 על מאיצי AI ובפרק 52 על נושא הדוקטורט שלו - מגבלות האצה של חומרה.היה זה אך טבעי שעדי יהיה המרואיין שיבוא לדבר איתנו על הסיבה שבגינה ההנפקה של סרברס הייתה מוצלחת כל כך.על מה דיברנו?- מה זה המוצר הזה של סרברס ואיזה בעיות הוא בא לפתור?- מה זה בכלל צ'יפ?- איך ואיפה מאחסנים מידע בצ'יפ כל כך ייחודי?- מה עושים עם הקירור?- איך מספקים מתח אחיד למפלצת ההספק הזו?- אילו אתגרים התמודדו איתם בסרברס בדרך לצ'יפ הזה?- מה הפתרונות האלטרנטיביים?- מה אנחנו מקווים שיקרה בעקבות ההנפקה הזו?אחרי שהאזנתם לפרק, ובהנחה שאתם לא בוטים, מוזמנים להצטרף לקבוצת הדיונים (השניה!) שלנו - שם תשמעו את הבום הבא לפני כולם >>> https://chat.whatsapp.com/BEGBOy60dNfAhWgc31vmXHנשמח לשמוע את דעתכם על הפרק בתגובות.פרק 95 - The Biggest ChipHard Reset - הפודקאסט של קהילת Hardware Engineering Israel.מוזמנים ליצור איתנו קשר במייל podcasthardreset@gmail.comהאזנה נעימה.

Bitcoiners - Live From Bitcoin Beach
Is Bitcoin Backed by Anything? How a Coffee Farm in El Salvador Proved Peter Schiff Wrong

Bitcoiners - Live From Bitcoin Beach

Play Episode Listen Later Jun 20, 2026 55:49 Transcription Available


Can Bitcoin save traditional farming from total collapse? El Salvador's Cherito (@CheritoCafe) de-financialized his supply chain, uses a Bitcoin circular economy to bypass banks, and runs nodes inside his coffee lab! In this episode, we sit down with Jorge, globally recognized in the community as Cherito, to dissect how legacy financial institutions manipulate commodity pricing to trap agricultural families in a perpetual cycle of poverty after leaving a lucrative technology and finance career in Seattle to return home to El Salvador. We took a look at how he is merging specialty coffee, agroecology, and a peer-to-peer Bitcoin standard to build an automated, zero-fiat circular economy from the ground up. Faced with total ruin from systemic economic decay and a devastating agricultural plague, Jorge realized that the only path to survival was to completely opt out of the generic commodity market by fundamentally restructuring how high-end specialty coffee is valued. Through the launch of the Bio Crop Project, he pioneered a localized framework that prioritizes biological resilience and ancestral agroecology over industrial chemicals and corporate-controlled seed monopolies, proving that wealth creation requires physical sacrifice, strategic crop diversification, and an unyielding commitment to local food independence.To fully protect this model from predatory global inflation, Jorge took absolute control of his supply chain by vertically integrating everything from the volcanic soil to direct international exporting. He breaks down how shielding delicate bean profiles from corporate exploitation relies entirely on understanding how to leverage the distinct microclimate and high-altitude terroir of El Salvador's mountains. By taking his message directly to global academic panels and roaster guilds, he transformed a vulnerable, localized crop into a premium global asset capable of dictating its own fair valuation on international markets.The structural revolution took place when this physical infrastructure collided with a digital, sound money standard right here in El Zonte. Jorge explains how finding the Bitcoin whitepaper provided the missing link for his business, allowing him to establish a completely self-sustaining circular economy that operates entirely outside the legacy banking grid. By anchoring his operations to a peer-to-peer electronic cash network, he successfully eliminated credit card processing traps, currency devaluation risks, and international wire delays, allowing his business to reallocate bank fees directly into the pockets of his local workers.We wrap up the conversation inside his newly launched flagship facility, where the boundaries of a traditional business are being completely redrawn through onsite coffee roasting, active ASIC miners, and a live blockchain node visualizer. We also dive into the fascinating history of his beautifully restored 1974 GM Cherito truck, which is a rare piece of Salvadorean automotive engineering that embodies absolute scarcity. Finally, Jorge reveals the genius behind upcycling discarded fruit skins into a refreshing, antioxidant-rich beverage the community now proudly calls "Bitcoin Libre", proving that under a true proof-of-work standard, absolutely nothing goes to waste.—Bitcoin Beach TeamLearn more about Cherito:X: https://x.com/CheritoCafeWeb: https://www.cheritocafe.com/IG: https://www.instagram.com/cheritocafe.svSupport and follow Bitcoin Beach:X: https://www.twitter.com/BitcoinBeachIG: https://www.instagram.com/bitcoinbeach_sv TikTok: https://www.tiktok.com/@livefrombitcoinbeach Web: https://www.bitcoinbeach.com Browse through this quick guide to learn more about the episode:00:00 Intro01:25 Why do tech professionals quit fiat careers for El Salvador?05:09 How does legacy financial structure impact farming stability?09:01 How do you bypass banking intermediaries via a circular economy?10:52 Why does physical proof of work require long-term capital?17:09 How does El Salvador protect commodities from global inflation?25:31 Why is specialty coffee cultivation analog proof of work?32:00 What does the bitcoin supply cap teach us about scarcity?45:24 How do you run a routing node and ASIC miners inside a business?49:24 Why is Bitcoin Libre a direct challenge to central planning?Live From Bitcoin Beach

Stephan Livera Podcast
The Evolution of Bitcoin Mining with NG Zhang SLP742

Stephan Livera Podcast

Play Episode Listen Later Jun 17, 2026 48:36


In this episode, NG Zhang, founder and CEO of Canaan Inc., shares his journey in Bitcoin mining technology, from early ASIC development to modern innovations like heat reuse and energy integration. We talk about the evolution of mining hardware, industry trends, and future prospects.Timestamps:(00:00) –  Intro and Early Days of Bitcoin Mining(03:40) –  From FPGA to ASIC Technology(06:22) –  The Avalon A16 Series(09:07) –  Reliability and Durability of Mining Machines(12:15) –  Failure Rates and Useful Life(14:40) - Primary Market and Secondary Market for Mining(16:40) –  Small Scale vs. Large Scale Mining(22:00) - What % is home mining today?(23:00) - Heat Recovery(29:00) - Different methods of cooling in Bitcoin Mining(31:40) - Mining machine form factor(32:40) - Stratum V2 Support?(37:35) - Market Dynamics: Public vs. Private Miners(39:50) –  Energy Grid Integration and Bitcoin Mining(44:25) - Efficiency gains are slowing down in bitcoin miningLinks: https://www.canaan.io/Stephan Livera links:Follow me on X: @stephanliveraSubscribe to the podcastSubscribe to Substack

The Crypto Conversation
Mitrade – Trading the Great Convergence

The Crypto Conversation

Play Episode Listen Later Jun 11, 2026 29:10


Cam Darlington is Global Strategy Expert at Mitrade, the Australian-founded CFD trading platform regulated by ASIC in Australia and CySEC in Europe, offering forex, commodities, indices, shares, and crypto from a single account. A Nova Scotia native based in Hong Kong for the past eight years, Cam brings a dual perspective: a traditional finance background working with brokers expanding across Asia, and recent experience as co-founder and COO of easy.fun, a social trading app built on Solana and Hyperliquid. Why you should listen Cam has a name for the phenomenon most people inside traditional brokerages never see from the trenches: the convergence. TradFi and crypto are collapsing into a single market structure, and the pivot point, he argues, is Washington. With the CLARITY Act working through the Senate and President Trump signing the Integrating Financial Technology Innovation into Regulatory Frameworks executive order in May, digital asset brokers are being ushered toward the core plumbing of the US financial system, including direct access to Federal Reserve payment rails. Add growing regulatory comfort with tokenized stocks trading at parity with their underlying assets, and the discount problem that dogged early real-world-asset experiments like Robinhood's tokenized equities starts to disappear. Tokenized RWAs, Cam says, just became viable. The second-order effects are reshaping market infrastructure itself. When a broker like Robinhood can mint tokenized stocks on its own proprietary chain and handle execution and settlement in-house, it stops feeding liquidity to the public exchanges. Cam frames the exchange's recent moves, including its tokenization partnership with Kraken built on the xStocks framework, as a defensive response to exactly this threat. He speaks from experience here: his team at easy.fun integrated the xStocks API and saw firsthand how thin liquidity gets once you trade beyond Nvidia, Apple, and Tesla. Cam says players most at risk are the centralized crypto exchanges, squeezed between newly crypto-enabled traditional brokerages on one side and purpose-built DeFi venues like Hyperliquid on the other. For traders, Cam's message is about survival. The first year determines whether someone becomes a trader or a statistic, and he is scathing about platforms offering 1,000x leverage to beginners, which he likens to handing a brand-new driver a Ferrari and pointing at the motorway. He makes the case for starting on a regulated platform with guardrails, modest leverage, built-in TradingView charting, and daily strategy feeds, which is precisely the gap Mitrade aims to fill as a companion to a traditional brokerage account.  Supporting links Stabull Finance Mitrade Sign up to Mitrade Andy on Twitter Brave New Coin on Twitter Brave New Coin If you enjoyed the show please subscribe to the Crypto Conversation and give us a 5-star rating and a positive review in whatever podcast app you are using.