POPULARITY
In this episode, Ray Cochrane breaks down Hot Chips 2026, the engineering conference where IBM, NVIDIA, Intel, AMD, Arm, and Fujitsu all showed how their next processors actually work. The headline disclosure is a mainframe core that runs Arm natively. Ray also covers Apple’s odd M6 Mac mini naming, London’s first autonomous Uber rides, Amazon’s purchase of the company behind DuckDB, GitHub’s HydraFusion, the best of IFA 2026, and new USDA research on farmed salmon. – Want to start a podcast? It’s easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens the show with a personal update. He apologizes for the late rollout and the missed Monday episode, having spent the week fighting a cold. He’s also heading to Michigan to spend time with family and visit his father’s gravesite. He hopes to record a couple of shows from his dad’s old studio while he is there, including a special episode planned for Tuesday. He also points listeners to a fresh site redesign that trades the old techie look for something cleaner and friendlier. The featured segment starts from a wrap-up post on Arm’s newsroom. However, Cochrane broadens it to cover the whole conference rather than a single article. Hot Chips has run every August since 1989, and this year’s event was the 38th, held August 23rd through the 25th at Stanford’s Memorial Auditorium. What Hot Chips Actually Is Cochrane draws a line between Hot Chips and the big consumer trade shows. CES and Computex exist for product announcements and marketing. Meanwhile, Hot Chips is an IEEE engineering conference where chip architects present block diagrams and die photos for thirty minutes at a stretch. The audience matters as much as the content. Roughly five hundred people who design chips for a living fill the room, and they would spot a fudged number immediately. No written paper is required, just the talk and the slides. In-person tickets sold out this year, as did Stanford’s dorm housing. Cochrane says he plans to cover the conference annually going forward. IBM Built a Mainframe Core That Speaks Arm The disclosure that stopped Cochrane cold came from IBM. Its next processor for IBM Z and LinuxONE runs two completely different instruction sets natively, on every one of its eleven cores. Those are z/Architecture, IBM’s own mainframe language, and AArch64, which is 64-bit Arm. Crucially, this is not emulation. Nor is it Arm cores glued onto the die beside the mainframe cores. IBM built 2,792 Arm instructions directly into the hardware, which it says is more than double the mainframe instruction count. Each core carries two separate decoders while sharing the caches, branch predictor and register files downstream. It switches between the two in nanoseconds. IBM even added dedicated hardware to flip byte order, because Arm and the mainframe store numbers in opposite directions. The payoff is Arm SystemReady compliance, meaning off-the-shelf Arm Linux runs on a mainframe unmodified. Patrick Kennedy of ServeTheHome, who was in the room, wrote: “I am sitting here still in awe of what IBM is doing here; this is not Z plus Arm cores, this is Z and Arm in one core.” Cochrane flags one precision point that is easy to get backward. IBM did not license Arm’s core designs and drop them in. Instead, it took its own mainframe core and taught it AArch64 under an architecture license, which is considerably harder engineering. The specifications are striking. The chip uses a 2nm process, with eleven cores running above 5.7GHz sustained and no turbo mode at all. Each core gets 36MB of L2 cache, backed by a 3.5GB virtual L4 pool. Furthermore, the reliability target is eight nines, which works out to roughly three tenths of one second of unplanned downtime per year. IBM gave it no name and no ship date, though the press expects “Telum III” around 2028. Arm, Fujitsu and NVIDIA Show Their Hands Arm itself had plenty to discuss, starting with an unfortunate name. Its first chip in thirty-five years is called the AGI CPU, which is a product name rather than any claim about artificial general intelligence. For three and a half decades, Arm designed processor blueprints and licensed them out, collecting royalties without competing. That era is now over. The AGI CPU is Arm’s own silicon, co-designed with lead customer Meta, running up to 136 cores on TSMC’s 3nm process at 300 watts. Arm’s CEO says the company has more than $2 billion in customer demand across the next two fiscal years. Fujitsu brought the detail Cochrane called the coolest of the conference. Its MONAKA chip packs 144 Arm-based cores, but the trick is the cache. Rather than sitting alongside the cores and eating die area, the entire last-level cache lives on a separate 5nm die with the 2nm compute die stacked directly on top. It ships in 2027 in 350-watt and 500-watt versions. NVIDIA had more stage time than anyone, with six sessions. Its new Vera CPU carries 88 cores of NVIDIA’s own Olympus design, which marks a change: the previous Grace CPU used Arm’s off-the-shelf cores. Consequently, Arm’s win here is the instruction set, not the blueprint. The memory disclosure drew the most attention, with a fully loaded system reaching 1.5TB at 1.2TB/s while the whole memory subsystem draws just 30 to 40 watts. The Caveat on NVIDIA’s Benchmark Slides Cochrane pushes back on how NVIDIA presented its numbers. On the standard SPEC integer benchmark, Vera scored 925 against AMD’s 128-core EPYC score of 898, about three percent ahead. However, the slide NVIDIA showed normalizes that same result per physical core, which makes a three percent gap look enormous. NVIDIA defends the choice, arguing that per-core throughput matters when thousands of AI agents run at once. Cochrane grants that it is a fair argument to make. Even so, his verdict is blunt: it is a different number from the headline one, and presenting it that way is not the best look. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. Apple’s Newest Chip Landed in Its Cheapest Mac Last week’s episode covered the Mac Studio half of Apple’s August announcement. Tonight Cochrane takes the other half, the new Mac mini, and finds the numbering genuinely strange. The $899 base Mac mini gets the M6, which is Apple’s first 2nm chip and the newest process the company has shipped. It brings twelve CPU cores, twelve GPU cores, and 170GB/s of memory bandwidth. Apple also introduced a third CPU core class called super cores. So Apple’s most advanced chip sits in its cheapest desktop, while the M5 Pro, M5 Max and M5 Ultra above it all carry a lower number. It gets stranger. The M6 mini has Thunderbolt 4 while the pricier M5 Pro mini has Thunderbolt 5, and the memory ceiling runs backward too. Apple explains none of it across three announcement pages. Reading between the lines, Cochrane figures the M6 is the entry point of a new generation that shipped ahead of its larger siblings. London’s Robotaxis Started Carrying Passengers Arm’s monthly roundup covers everything outside the data center, and Wayve stood out. Transport for London granted the British self-driving company private hire vehicle licenses on August 5th, the same category a minicab needs. Then on September 3rd the service launched with Uber, marking the first autonomous rides ever offered to UK passengers. A small fleet of Ford Mustang Mach-Es covers anywhere in London except the airports, with a TfL-licensed safety driver still aboard. Over 140,000 Londoners signed up. The compute runs on NVIDIA’s Arm-based automotive platform, which is why Arm claims the win. Cochrane notes the pattern: once you set a standard nobody can move off, these wins keep arriving. Two more items round it out, both about squeezing AI onto phones. Google’s Pixel 11 shipped in August with the Tensor G6, and Google claims on-device AI runs up to 3.5x faster using 3.5x less energy. Those are Google’s own unbenchmarked figures. Separately, Graphcore’s research team worked with Arm to run an 11-billion-parameter vision model on phone-class processors by squeezing each parameter to 2.7 bits, taking the model from roughly 22GB down to 3.7GB. Intel Is Pitching AI Infrastructure From a Long Way Back Intel’s newsroom post previews the AI Infra Summit, running September 15th through 17th in Santa Clara. CEO Lip-Bu Tan takes a fireside chat on Tuesday morning, with three Intel sessions across the show. Cochrane unpacks two terms first. Physical AI means AI that acts in the real world through sensors and motors rather than living on a screen, and Intel takes it seriously enough to have renamed its PC division the Client Computing and Physical AI Group in May. Disaggregated inference splits the two phases of running a model: reading your prompt is compute-hungry, while writing the answer back is memory-hungry. The context is where this gets interesting. Intel’s revenue rose 25 percent last quarter, its fastest growth since 2011. Nevertheless, in AI accelerators the company barely registers next to NVIDIA. Gaudi has effectively been abandoned, to the point that Intel stopped maintaining its open-source driver, and AMD passed Intel in data center revenue last quarter. Tellingly, when Intel demoed disaggregated inference at Computex, NVIDIA GPUs handled one phase and SambaNova chips the other while Intel supplied the coordinating CPU. Crescent Island, Intel’s actual inference chip, does not sample until later this year, and Intel declined to publish its memory bandwidth. Amazon Bought the Company Behind DuckDB On August 26th, Amazon signed a deal to acquire DuckLabs, the Amsterdam company behind DuckDB. Cochrane spends time explaining what DuckDB is, since listeners outside the data world may never have encountered it. The problem it solves is familiar. Querying a large pile of data files traditionally meant either running a database server or spinning up a data warehouse with a cluster, a bill, and a loading pipeline. Both are heavy machinery for a question you wanted answered in ten seconds. DuckDB instead ships as a library rather than a server. You add it to your program like any other package, point it at your files, and write ordinary SQL directly against them. The data never moves. The usual comparison is SQLite, which is embedded in nearly every phone and browser on earth. Where SQLite excels at looking up one record, DuckDB rebuilds that embedded idea for chewing through millions of rows. It now sees roughly 62 million monthly downloads on Python’s package index alone, up from about 25 million last October. AWS says it is buying the company, not the project. DuckDB stays free and open source under the MIT license, held by a Dutch nonprofit foundation, and the founders join AWS while continuing to run technical direction from Amsterdam. Andy Warfield, a VP and distinguished engineer at AWS, described DuckDB as “the glibc of structured data: a lean, unglamorous, ubiquitous dependency that a great deal of software links against and almost nobody has to think about.” Cochrane sits with what that ownership means. He reaches for an analogy: imagine Daniel Stenberg selling curl. He doubts it would ever happen, but the concept alone is startling given how much infrastructure depends on it. His read is that AWS is betting DuckDB becomes as foundational as curl and SQLite already are. GitHub Has One AI Model Grade Another One’s Homework GitHub shipped Project HydraFusion into Copilot as a research preview. Instead of routing your request to a single model, it picks one of three approaches per request. Sometimes one model simply answers. Alternatively, a cheaper model drafts, and a quality gate decides whether to escalate. The interesting one is Critique. One model writes the code, a separate model from a different family reviews it read-only, and the original gets one pass to revise. Despite the name, nothing is fused here. There is no voting and no merging, just one model at a time with a gate deciding whether to spend more. GitHub explained the reasoning in an earlier post: “a model reviewing its own work is still bounded by its own training biases: the same training data and techniques, the same blind spots.” Research supports it. A team at NeurIPS in 2024 showed that models recognize their own writing and score it higher than human graders do. GitHub’s earlier number had a Claude Sonnet and GPT critic pairing closing about three-quarters of the gap between Sonnet and the larger Opus model. This mirrors a workflow Cochrane uses constantly and has described on a previous episode. He runs a cross-check review with a second model from a different company, and it routinely surfaces issues the first model missed. He explains that different training data, different engineers, and different reinforcement approaches build different internal biases about what counts as correct. Looking ahead, he expects more of these “Frankenstein patterns” where models from different training families work together. The Best of IFA 2026, and What You Can Actually Buy IFA opened to the public in Berlin for its 102nd year, with about 1,900 brands. Cochrane splits The Verge’s roundup in two, since much of what generates headlines at these shows never ships. Starting with real products, iRobot’s flagship Roomba Max 875 Combo runs $1,199 and ships in about two weeks. Its SealForce feature drops a hidden skirt from the chassis when it detects carpet, sealing against the fibers so suction concentrates instead of leaking out the sides. That reaches 35,000 pascals, iRobot’s strongest yet. A step-down model at $899 carries the same trick, though Cochrane balks at both prices. Anker’s Soundcore Sleep 4 Pro earbuds arrive in November at $349.99. The charging case carries its own round touchscreen, so you pick soundscapes, set alarms, and read sleep stats without your phone. Optical sensors read heart rate and variability from the ear canal, which beats the wrist for accuracy, and the case masks a snoring partner. Cochrane remains unconvinced about sleeping with earbuds in. Philips also has smart rope lights, the Hue Liane 360, on sale now. They glow evenly around the tube rather than showing individual LEDs. They also run $400 for three meters, which works out to about $130 per meter of rope light. As for concepts nobody can buy, iRobot showed a robot vacuum that carries a smaller robot vacuum on its back in a garage and lowers it to deploy. Lenovo brought a 14-inch laptop whose screen rolls out to 17 inches at the press of a button, which reviewers call the first rollable that feels close to shippable. Tecno showed a phone with essentially no border around the screen, and Acer had a Windows gaming handheld that swivels its screen up over a keyboard. Two themes ran through the show. Humanoid robots were the loudest thing on the floor, and IFA’s own CEO framed the event as being about robots that work rather than robots that demo. Meanwhile, AI stopped being its own product category and became an ingredient, showing up in refrigerators, treadmills, dishwashers, and motorized TV mounts. Farmed Salmon Isn’t the Omega-3 Machine It Used to Be USDA scientists measured farmed salmon and found considerably less of the good fat than the government’s own database claims. EPA and DHA are two fatty acids you get almost entirely from fish. Your body can build them from the plant version, but only in tiny amounts, so the NIH’s position is that eating them is the only practical way to raise your levels. Those fatty acids are structural pieces of every cell, with DHA concentrating in the brain and retina. That is why the federal dietary guidelines, issued jointly by USDA and Health and Human Services, recommend at least eight ounces of fish a week and steer you toward salmon. That amount is calibrated to deliver about 250 milligrams a day. Published in Frontiers in Nutrition last month, the study found EPA and DHA in farmed Atlantic salmon came in 54.7 percent lower than USDA’s own reference values, last updated in 2018. A three-ounce serving fell from roughly 1,670 milligrams to about 756. Consequently, two servings a week now fall about 14 percent short of the target. Plant-derived fats meanwhile rose two to three times over. The likely cause is feed. Salmon are carnivores, and farms once fed them oily little fish. There was never going to be enough of those as the industry scaled, so crops filled the gap: soy, canola, sunflower and linseed. Importantly, the study does not claim to have proven this and calls the feed shift a plausible explanation. Independent corroboration lends it credibility. Researchers at Stirling measured a similar halving in Scottish salmon between 2006 and 2015, and Norway’s marine institute saw it across thousands of samples. There is a land dimension too. Roughly half the world’s soy grows in South America, where rainforest gets cleared for feed. Matthew Hayek, who studies the environmental cost of protein at NYU, told Inside Climate News that “soy is a major, important protein and oil ingredient in fish farming.” Adding up two decades of soy across all fish farming, he puts the extra forest clearing at around the area of Nicaragua or Bangladesh. That figure covers all fish farming rather than salmon alone, and Hayek notes it is hard to attribute soy use to any single species. Cochrane closes with two caveats. First, the study measured only eight fish, bought around Maryland, DC and Virginia over six weeks in 2023, and nearly all sourced from Chile. That is not a national survey, and the authors say plainly the sample was not large enough to change government advice. Rather, it flags that a federal database value needs rechecking, which is what the paper set out to do. Second, on whether you should care, farmed salmon still beats beef, chicken and eggs by a mile, since those carry essentially zero EPA and DHA. What it loses is its crown among fatty fish, dropping to mid-pack behind herring, sardines and mackerel and roughly level with trout. The broader health case is also softer than the 2000s suggested. A review of 86 trials covering 162,000 people found supplements barely moved heart attacks or deaths, so eating fish and swallowing fish oil are not the same claim. No producer has responded to the findings, and USDA, whose own scientists ran the study, declined an interview and did not answer emailed questions. Cochrane wraps up with housekeeping and a note that he will be back on Labor Day. The post The Mainframe Learned to Speak Arm #1875 appeared first on Geek News Central.
Walmart collected about $2.9 billion in tariff refunds and spent it on roughly 11,000 rollbacks in Walmart US. Transactions grew and operating income rose 28.8%. The comp still slowed to 2.6% excluding fuel, the weakest quarter since 2020, with the softness concentrated in lower-income households. Walmart raised full-year guidance on the assumption that the second half improves on the back of that price investment, which puts a refund that will not repeat into the base of next year's math.Lowe's earned $4.27 a share on $2 billion more revenue than last year, when it also earned $4.27. Comparable sales rose two tenths of one percent. Almost all of the revenue growth was acquired, from a building products distributor and an interior finishes installer that sell into new residential construction, and Lowe's removed the top of its full-year outlook four separate times on the call. Online grew 15.7%. The release blames persistent do-it-yourself macro pressure for the rest, which is a long way of saying the Saturday deck lumber customer has not come back.Target's traffic did come back. Comps grew 3.8% with 3.6 points from traffic, and apparel and accessories grew $4 million on a $4 billion base.Ipsy is launching a marketing services arm and will no longer say what its revenue is. Six years ago it published 4.3 million subscribers and a billion dollars.Plus the Investor Minute: Ferrero buys Purely Elizabeth, Amazon buys DuckDB Labs but not DuckDB, Mubadala takes majority control of Arrive Logistics, Blank Street raises $105 million, Lavanta raises $22 million.The Watson Weekly is sponsored by Avalara. Tax compliance gets harder with every new channel, state, product and market. See what Avalara Agentic Tax and Compliance does about it at avalara.watsonweekly.com#watsonweekly #walmart #lowes #target #ipsy
On this episode of The Joe Reis Show, I'm joined by Dan Bennett, Head of Technology for the Enterprise Data Organization at S&P Global.We get into what it actually looks like to build with modern AI coding tools like Claude Code. Not just as a toy, but for writing production-grade C++. Dan walks through how he built an open-source RDF extension for DuckDB, why rock-solid test coverage is non-negotiable when working with LLMs, and why engineering leaders need to keep their hands dirty to understand where this tech is headed. We also react to the breaking news of AWS acquiring DuckDB Labs, talk through the shift from "human-in-the-loop" to autonomous agents running in headless VMs, and dive into why data semantics across organizational boundaries remains one of the hardest - and most important - unsolved problems in our industry.Website: https://nonodename.com/
Building something people actually want is supposed to be the happy ending. But it arrives with a bill attached: feature requests you didn't ask for, pull requests you'd rather not maintain forever, users demanding the one thing you swore you'd never build, and — if you're unlucky — a company where the sales team quietly starts deciding what engineering works on. DuckDB has spent the last two and a half years working through that list. So how do you stay a database engineering team when success keeps trying to turn you into something else?Hannes Mühleisen, co-creator of DuckDB, is back to talk through the answers they've landed on. Their fix for unwanted pull requests was an extension mechanism, which then forced them to make every part of the engine pluggable — including the parser, which meant ripping out 20,000 lines of Postgres' yacc grammar and rewriting SQL parsing on top of PEG. Their fix for the users demanding client-server was Quack, a protocol designed by people who'd already published a paper on why every existing database wire protocol is wrong. And their answer to Apache Iceberg, after three years of implementing it, was DuckLake: throw out the Avro-and-JSON metadata files and keep the metadata in a database, on the grounds that the Iceberg REST catalog has a Postgres in it anyway.Which brings us to the news Hannes breaks in this episode: DuckDB Labs is being acquired by AWS, while the DuckDB Foundation, the project and its licence stay where they are. There's the question of why a profitable, self-funded, 30-person company in Amsterdam would take that deal, what commitments you write into the contracts when you're worried today's promises might outlive today's management, and what it's actually like to have a boss again after five years without one. If you're curious how an open source project keeps its technical soul once the enterprise arrives — or you just want to know why parsing SQL is harder than parsing almost anything else — Hannes has some good answers.---Support Developer Voices on Patreon: https://patreon.com/DeveloperVoicesSupport Developer Voices on YouTube: https://www.youtube.com/@DeveloperVoices/joinOur previous episode with Hannes: https://youtu.be/pZV9FvdKmLcDuckDB: https://duckdb.org/DuckDB Foundation: https://duckdb.foundation/DuckLabs (formerly DuckDB Labs): https://ducklabs.com/DuckLake: https://ducklake.select/Quack (DuckDB's client-server protocol): https://duckdb.org/quack/DuckDB v2.0: Your Database Deserves a Better Parser: https://duckdb.org/2026/08/20/duckdb-20-peg-parserRuntime-Extensible Parsers (CIDR 2025 paper): https://duckdb.org/pdf/CIDR2025-muehleisen-raasveldt-extensible-parsers.pdfDon't Hold My Data Hostage (VLDB 2017 paper): https://www.vldb.org/pvldb/vol10/p1022-muehleisen.pdfcpp-peglib: https://github.com/yhirose/cpp-peglibGNU Bison: https://www.gnu.org/software/bison/PEP 617 – New PEG parser for CPython: https://peps.python.org/pep-0617/PRQL: https://prql-lang.org/Apache Iceberg: https://iceberg.apache.org/CWI (Centrum Wiskunde & Informatica): https://www.cwi.nl/en/DuckCon #7, Amsterdam: https://duckdb.org/events/2026/06/24/duckcon7/Kris on Bluesky: https://bsky.app/profile/krisajenkins.bsky.socialKris on Mastodon: http://mastodon.social/@krisajenkinsKris on LinkedIn: https://www.linkedin.com/in/krisjenkins/
In this talk, Radovan Bacovic, Principal Data Engineer and Snowflake/dbt Ambassador, shares his two-decade journey, from traditional database administration to leading modern data platform engineering. We explore the essential building blocks of scalable data architectures, the trade-offs of no-code solutions, and how to effectively integrate AI into your data pipelines while maintaining strict security and governance.You'll learn about:- The evolution of the full-stack data engineer over the past two decades.- Building scalable, business-driven data platforms without premature optimization.- Efficient data modeling and transformation using DuckDB and dbt.- Leveraging AI to accelerate pipeline development and launch data products.- Implementing enterprise data governance and robust pipeline security.- Balancing no-code integration tools with strict DataOps methodologies.LINKS:https://gitlab.com/radovan.bacovic/wsc26_dataops/-/blob/main/README.md?ref_type=headsTIMECODES:0:00 Data Engineering Career Trajectory and Industry Experience6:37 Full Stack Data Engineer Role Evolution and Responsibilities14:45 Scalable Data Platform Architecture and Business Alignment21:59 DuckDB Integration and Premature Database Optimization Avoidance30:04 Data Pipeline Transformation and Data Modeling with dbt35:31 Modern Data Platform Ecosystem and Infrastructure Orchestration42:13 AI Powered Data Products and LLM Pipeline Automation49:42 Enterprise Data Governance and Pipeline Security Best Practices55:06 No Code Data Integration Tools and DataOps MethodologyThis talk is ideal for data engineers, analytics engineers, and data platform leads looking to modernize their tech stacks and future-proof their infrastructure. It provides highly actionable insights for anyone aiming to build resilient, AI-ready data platforms that prioritize real business value and strict security constraints.Connect with Radovan- Twitter - https://twitter.com/Al_Grigor - Linkedin - https://www.linkedin.com/in/agrigorev/ Connect with DataTalks.Club:- Join the community - https://datatalks.club/slack.html- Subscribe to our Google calendar to have all our events in your calendar - https://calendar.google.com/calendar/r?cid=ZjhxaWRqbnEwamhzY3A4ODA5azFlZ2hzNjBAZ3JvdXAuY2FsZW5kYXIuZ29vZ2xlLmNvbQ- Check other upcoming events - https://lu.ma/dtc-events- GitHub: https://github.com/DataTalksClub- LinkedIn - https://www.linkedin.com/company/datatalks-club/ - Twitter - https://twitter.com/DataTalksClub - Website - https://datatalks.club/
Marie-Cécile Riom est AI Product Specialist chez Snowflake, la plateforme Data & IA que tout le monde connaît. Snowflake connaît depuis des années une croissance exceptionnelle. De nombreuses entreprises telles que Qonto, Sanofi et Swile l'utilisent au quotidien.On aborde :
Jordan Tigani helped build BigQuery, then left to bet that most data isn't big. Three years on, agents are proving him right. The MotherDuck CEO joins Tristan Handy on why local-first databases fit the agent era, and what an "agent swarm for data management" looks like. For full show notes and to read the podcast's companion newsletter, head to https://roundup.getdbt.com. The Analytics Engineering Podcast is sponsored by dbt Labs.
This week the trio covers the Latest Ubuntu, Fedora, and CachyOS news. Btrfs has a big performance win, USB4 brings fast data transfers, the latest kernel RC has prompted a classic Torvalds rant. And then Jonathan flies in to wrap up the show with Open Source AI definition news. For tips, we have quein for turbo-charges who is, Shelly for smarter package management, htmlq for querying a web page, and DuckDB for slick SQL on the command line. You can find the show notes at https://bit.ly/434Hrkg and enjoy! Host: Jonathan Bennett Co-Hosts: Ken McDonald, Rob Campbell, and Jeff Massie Download or subscribe to Untitled Linux Show at https://twit.tv/shows/untitled-linux-show Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Club TWiT members can discuss this episode and leave feedback in the Club TWiT Discord.
This week the trio covers the Latest Ubuntu, Fedora, and CachyOS news. Btrfs has a big performance win, USB4 brings fast data transfers, the latest kernel RC has prompted a classic Torvalds rant. And then Jonathan flies in to wrap up the show with Open Source AI definition news. For tips, we have quein for turbo-charges who is, Shelly for smarter package management, htmlq for querying a web page, and DuckDB for slick SQL on the command line. You can find the show notes at https://bit.ly/434Hrkg and enjoy! Host: Jonathan Bennett Co-Hosts: Ken McDonald, Rob Campbell, and Jeff Massie Download or subscribe to Untitled Linux Show at https://twit.tv/shows/untitled-linux-show Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Club TWiT members can discuss this episode and leave feedback in the Club TWiT Discord.
It's been a few months on the road, bouncing through San Francisco a bunch, across Asia and Europe, and a quick stop in Detroit. In this audio-only Freestyle Friday I unpack what I've been seeing out there. If I had to pick one word for the mood worldwide, it's uncertainty: energy and supply shocks rippling out of the Middle East, fuel and resource shortages, flights getting canceled with no notice, and AI scrambling the playbook for vendors, practitioners, and leaders alike.I get into why so many data tooling companies are quietly having existential conversations, how Atlan tore its product down to rebuild AI-native (a full conversation with Prukalpa is coming next week), and a fun experiment I shipped this week with DuckDB Quack.I also dig into the split I keep seeing: senior practitioners getting superpowers while juniors face a brutal job market, leaders being asked to do far more with less, and why I think the industrial-age org chart is finally on its way out.Plus some personal updates: the new book is now targeting late July and a companion course is on the way.Finally, I'm mixing audio and video formats going forward (Freestyle Friday will probably be mostly audio), the Practical Data Community newsletter is live, and there's a Salt Lake City conference brewing for late January. Lots in the hopper...------------------This episode is sponsored by Revefi, who gives you full cost and performance visibility for Snowflake by warehouse, user, and workload. One team cut Snowflake costs ~50% across 711 warehouses in under 48 hours. Book a demo at revefi.com/demo.------------Timestamps0:00 — Intro & travel recap — Sets the stage: months of globe-trotting across Asia, Europe, and the US1:10 — Global uncertainty & resource scarcity — Fuel/water shortages in Southeast Asia, flight cancellations in Europe, ripple effects of geopolitical tensions5:30 — AI dominates every conversation — The #1 topic at conferences worldwide; vendors facing existential questions and forced to rethink everything (Atlan pivot, DuckDB agent idea)10:14 — AI's impact on workers at every level — Senior practitioners gaining superpowers, juniors worried about jobs, leaders expected to do more with less17:51 — Key takeaway: everyone feels behind — Even top AI insiders are uncertain; give yourself grace, upskill, and consider building something for yourself20:38 — Announcements — Book drops July 27th, course coming, Practical Data Community Newsletter live, fall travel schedule (London, Paris, possible Salt Lake City conference)
At a recent MCP developer summit, The New Stack spoke with Till Döhmen, AI lead atMotherDuck, about the company's growing role in the evolving DuckDB ecosystem. Backed by investors includingTomasz Tunguz, MotherDuck is commercializing the open-source analytical databaseDuckDBwhile also expanding how employees interact with data through AI agents rather than traditional dashboards. Döhmen emphasized the company's close collaboration withDuckDB FoundationandDuckDB Labs. Because MotherDuck operates what he described as the world's largest fleet of DuckDB databases, the startup regularly pushes the database to its limits and feeds insights back to the core maintainers. Rather than forking DuckDB to create proprietary advantages, MotherDuck instead extends the platform through its existing architecture while contributing core improvements upstream when needed. The conversation highlighted the delicate but productive relationship between venture-backed companies and the open-source projects they commercialize, positioning MotherDuck as another example of startups driving both OSS adoption and strong business growth simultaneously. Learn more from The New Stack around the latest in DuckDB: DuckDB: Query Processing Is King DuckDB: In-Process Python Analytics for Not-Quite-Big Data Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Index drauf und fertig. Klingt nach einem soliden Plan, oder? Leider nur so lange, bis die Daten wachsen, der Workload kippt oder der Optimizer plötzlich andere Entscheidungen trifft. Dann wird aus dem vermeintlichen Performance-Booster schnell ein Bremsklotz. Genau hier steigen wir in dieser Episode ein und schauen uns an, warum Indexstrukturen in Datenbanken viel mehr sind als ein technischer Quick Fix.Wir sprechen darüber, was ein Index eigentlich ist, wie Datenstruktur, Algorithmus, Hardware und Workload zusammenhängen und warum Begriffe wie Selektivität, Kardinalität, Full Table Scan, Write Amplification und Cache-Lokalität in der Praxis entscheidend sind. Außerdem schauen wir auf typische Datenbank-Themen wie Primary Key, B-Tree, Binary Search, Covering Index, Optimizer, Slow Query Log und Explain Statements. Dabei wird auch klar, warum ein Index manchmal hilft, manchmal ignoriert wird und manchmal sogar langsamer ist als gar kein Index.Wenn du mit PostgreSQL, MySQL, MariaDB oder ganz allgemein mit Datenbank-Performance arbeitest, bekommst du hier ein solides Fundament und einige praktische Denkanstöße für deinen Alltag als Softwareentwickler:in. Und ja, wir sprechen auch über Invisible Indexes in MySQL. Ein Feature, das fast wie ein Zaubertrick klingt, aber beim Testen und beim sicheren Aufräumen von Legacy-Systemen überraschend praktisch sein kann. Viel Spaß beim Hören und vielleicht beim anschließenden Blick auf dein Datenbankschema.Unsere aktuellen Werbepartner findest du auf https://engineeringkiosk.dev/partnersDas schnelle Feedback zur Episode:
Josh Wills has spent 25 years writing data pipelines, with a career spanning Cloudera, as Director of Data Engineering at Slack, on the dbt DuckDB adapter, and now training foundation models at Datology AI. He uses coding agents every day. And he keeps running into the same wall: the agents jump to conclusions, fix the wrong thing, and ship pipelines no one understands.In this conversation, we unpack why AI agents struggle with the messiest, highest-stakes parts of data work, and what it means for the engineers managing them.We get into:- Big Data is back- Why AI agents jump to conclusions on benchmarks and complex bottlenecks- The $200K vibe-coded pipeline problem nobody wants to talk about- Why there's no training data for the gnarly enterprise pipelines that actually power businesses- "We're all managers now" - managing unreliable agents like managing unreliable people- Wicked problems and the limits of intelligence- Why politics is the last human endeavor to fall to LLMs (the data is never written down)- Whether classical ML still has a place (yes)- What Josh would tell a new grad starting in data today
Java und Performance in einem Satz? Für viele klingt das immer noch wie ein Widerspruch. Dann kommt eine Challenge daher, bei der eine Milliarde Zeilen Wetterdaten verarbeitet werden sollen, und plötzlich wird aus Stammtischwissen ein echter Engineering-Nerdfight. Genau darum geht es in dieser Episode. Wir tauchen tief in die One Billion Row Challenge ein und schauen uns an, wie eine vermeintlich einfache Aufgabe zum internationalen Performance-Contest wurde.Wir sprechen darüber, warum Gunnar Morling diese Challenge gestartet hat, wie aus einer naiven Lösung mit fast fünf Minuten Laufzeit optimierte Implementierungen mit rund 1,5 Sekunden wurden und welche Rolle dabei Java, GraalVM, Memory Mapping, Unsafe, SIMD, Branchless Coding, Hashmaps, Cache-Lines und Integer-Arithmetik spielen. Außerdem schauen wir auf die Kritik an der Challenge, etwa RAM-Disk, Dataset-Overfitting und CPU-spezifische Optimierungen, und wir werfen einen Blick auf alternative Umsetzungen in C, Go, PHP, SQL, DuckDB, ClickHouse, AWK und sogar auf GPU-Ansätze.Wenn du Performance-Optimierung nicht nur als Buzzword, sondern als Mischung aus Hardware-Verständnis, Datenstrukturen, Compiler-Wissen und Community-Lernen sehen willst, bist du hier genau richtig. Und ganz nebenbei klären wir auch noch, ob Java wirklich langsam ist oder ob dieser Mythos endlich in Rente darf.Bonus: AWK schafft es in elf Zeilen. Nicht schnell, aber stilvoll.Unsere aktuellen Werbepartner findest du auf https://engineeringkiosk.dev/partnersDas schnelle Feedback zur Episode:
AI has completely inverted how we build and scale software, which begs the question: What exactly is a moat anymore? In this Freestyle Friday, recovering from jet lag and hiking through the beautiful hills of Salt Lake City, I'm breaking down a recent conversation with a VC friend about defensibility in the era of coding agents. I also look at this through Charlie Munger's lens of "inversion" to figure out what isn't a moat anymore (spoiler: thin foundation model wrappers, "AI", and feature velocity are dead).I also dive into what is defensible today, from mission-critical systems of record like DuckDB and Postgres, to personal branding, to shifting SaaS pricing from per-seat to per-token.
This interview was recorded for GOTO Unscripted.https://gotopia.techCheck out more here:https://gotopia.tech/articles/421Félix GV - Current Interests: Multi-Planetary Databases, Data Sovereignty & LifeloggingOlimpiu Pop - Technologist & Tech JournalistRESOURCESFélixhttps://bsky.app/profile/felixgv.ninjahttps://github.com/FelixGVhttps://www.linkedin.com/in/felixgvOlimpiuhttps://x.com/olimpiupophttps://github.com/zrollhttps://www.linkedin.com/in/olimpiupopLinkshttps://venicedb.orghttps://github.com/linkedin/venicehttps://rocksdb.orghttps://duckdb.orgDESCRIPTIONFélix GV, a former engineer at LinkedIn and architect of the Venice database system, discusses the complexity of building planetary-scale data systems. He explains Venice's unbundled architecture where each component—from Kafka-based pub/sub to RocksDB-powered servers—operates as an independent distributed system. Félix details their rigorous chaos engineering practices, including regular load tests that push data centers beyond normal capacity to ensure reliability.The discussion covers fundamental distributed systems concepts like the CAP theorem and the trade-offs between consistency and availability in multi-region deployments. He also explains why Venice, as a derived data system, deliberately sacrifices strong consistency for high throughput and availability, and concludes by discussing their experimental integration of DuckDB for SQL-based analytics and data exploration capabilities.RECOMMENDED BOOKSKasun Indrasiri & Danesh Kuruppu • gRPC: Up and Running • https://amzn.to/3sBGBJJTomer Shiran, Jason Hughes & Alex Merced • Apache Iceberg: The Definitive Guide • https://amzn.to/488Z30kWilliam Smith • Arrow Flight Protocols and Practices • https://amzn.to/4o2Q2fdAdi Polak • Scaling Machine Learning with Spark • https://amzn.to/3N9vx1HMark Needham, Michael Hunger & Michael Simons • DuckDB in Action • https://amzn.to/45QwSliSimon Aubury & Ned Letcher • Getting Started with DuckDB • https://amzn.to/3VPk4qBlueskyInstagramLinkedInFacebookCHANNEL MEMBERSHIP BONUSJoin this channel to get early access to videos & other perks:https://www.youtube.com/channel/UCs_tLP3AiwYKwdUHpltJPuA/joinLooking for a unique learning experience?Attend the next GOTO conference near you! Get your ticket: gotopia.techSUBSCRIBE TO OUR YOUTUBE CHANNEL - new videos posted daily!
Vincent Heuschling reçoit Hayssam Saleh, créateur de **Starlake**, une plateforme data open source française née de la factorisation de projets clients depuis 2017-2018. L'épisode intervient dans un contexte de consolidation du marché (rachat de DBT et de SQLMesh par Fivetran), qui invite à challenger les solutions établies.Starlake se distingue par une approche **entièrement déclarative** (YAML + SQL natif, sans Jinja) couvrant toute la chaîne data engineering : ingestion, transformation, orchestration et qualité des données. L'outil s'appuie sur les moteurs sous-jacents des plateformes cibles (Snowflake, BigQuery, Spark) et génère automatiquement les DAGs pour les orchestrateurs du marché (Airflow, Dagster, Snowflake Tasks).Parmi les fonctionnalités marquantes : le **data branching** (branches de données à la manière de Git), l'inférence automatique de schémas YAML à partir de fichiers sources, un **transpiler SQL** multi-plateformes, et l'extraction du lineage depuis du SQL brut sans annotation. L'intégration récente de **DuckLake** ouvre la voie à des architectures on-premise souveraines à coût maîtrisé (sous 300 €/mois sur OVH, Scaleway, Clever Cloud).Le modèle économique repose sur le support, la formation, et le consulting : Starlake s'installe dans le cloud du client, avec mise à jour automatique gérée par l'équipe, sans accès aux données.**Chapitres****00:00:27** – Introduction : consolidation du marché data (rachat de DBT et SQLMesh par Fivetran) et présentation de l'épisode**00:03:13** – Hayssam et la genèse de Starlake : parcours Spark/Scala, POC à 4 000 formats de fichiers (2017-2018)**00:09:51** – Architecture et philosophie : load, transform, orchestration unifiés en déclaratif (YAML + SQL natif, pas de Jinja)**00:00:18:18** – Starlake vs DBT : différences philosophiques, composabilité, fonctionnalités 100 % open source**00:00:22:20** – Data branching, Starlake Labs (pipe syntax, transpiler SQL, lineage) et expérience développeur (DuckDB local, UI point-and-click)**00:36:35** – Modèle open source et économique : licence Apache, support, formation, marketplace cloud souveraine**00:43:42** – DuckLake : alternative on-premise/cloud souverain (OVH, Scaleway, Clever Cloud) et comment contribuer / démarrer**Le BigdataHebdo**Le BigdataHebdo est le podcast Francophone de la Data et de l'IA.Retrouvez plus de 200 épisodes https://bigdatahebdo.comRejoignez la communauté sur le Slack https://join.slack.com/t/bigdatahebdo/shared_invite/zt-a931fdhj-8ICbl9dbsZZbTcze61rr~Q
Kennst du diese Situation im Team: Jemand sagt "das skaliert nicht", und plötzlich steht der Datenbankwechsel schneller im Raum als die eigentliche Frage nach dem Warum? Genau da packen wir an. Denn in vielen Systemen entscheidet nicht das nächste hippe Tool von Hacker News, sondern etwas viel Grundsätzlicheres: Datenlayout und Zugriffsmuster.In dieser Episode gehen wir einmal tief runter in den Storage-Stack. Wir schauen uns an, warum Row-Oriented-Datastores der Standard für klassische OLTP-Workloads sind und warum "SELECT id" trotzdem oft fast genauso teuer ist wie "SELECT *". Danach drehen wir die Tabelle um 90 Grad: Column Stores für OLAP, Aggregationen über viele Zeilen, Spalten-Pruning, Kompression, SIMD und warum ClickHouse, BigQuery, Snowflake oder Redshift bei Analytics so absurd schnell werden können.Und dann wird es file-basiert: CSV bekommt sein verdientes Fett weg, Apache Parquet seinen Hype, inklusive Row Groups, Metadaten im Footer und warum das für Streaming und Object Storage so gut passt. Mit Apache Iceberg setzen wir noch eine Management-Schicht oben drauf: Snapshots, Time Travel, paralleles Schreiben und das ganze Data-Lake-Feeling. Zum Schluss landen wir da, wo es richtig weh tut, beziehungsweise richtig Geld spart: Storage und Compute trennen, Tiered Storage, Kafka Connect bis Prometheus und Observability-Kosten.Wenn du beim nächsten "das skaliert nicht" nicht direkt die Datenbank tauschen willst, sondern erst mal die richtigen Fragen stellen möchtest, ist das deine Folge.Bonus: DuckDB als kleines Taschenmesser für CSV, JSON und SQL kann dein nächstes Wochenend-Experiment werden.Unsere aktuellen Werbepartner findest du auf https://engineeringkiosk.dev/partnersDas schnelle Feedback zur Episode:
In episode 31 of Open Source Ready, Brian and John sit down with Matthaus Krzykowski, Thierry Jean, and Elvis Kahoro to explore how dlt and dltHub are changing the way developers build data pipelines. The conversation dives into DuckDB, LLM-driven workflows, and the growing shift toward developer-first data engineering. They also discuss open source adoption, AI orchestration, and what it means to be a “10x engineer” in 2026.
This interview was recorded for GOTO Unscripted.https://gotopia.techCheck out more here:https://gotopia.tech/articles/412Andrew Lamb - Staff Engineer at InfluxData, ASF Member & PMC Apache DataFusion & Apache ArrowOlimpiu Pop - Technologist & Tech JournalistRESOURCESAndrewhttps://bsky.app/profile/andrewlamb1111.bsky.socialhttps://x.com/andrewlamb1111https://github.com/alambhttps://www.linkedin.com/in/andrewalambhttps://andrew.nerdnetworks.orgOlimpiuhttps://x.com/olimpiupophttps://github.com/zrollhttps://www.linkedin.com/in/olimpiupopLinkshttps://www.influxdata.com/blog/flight-datafusion-arrow-parquet-fdap-architecture-influxdbhttps://www.cidrdb.org/cidr2005/papers/P19.pdfDESCRIPTIONOlimpiu Pop speaks with Andrew Lamb, staff engineer at InfluxData and PMC member of Apache DataFusion and Apache Arrow, about how modern data systems are built using standardized open source components rather than being developed from scratch.Andrew discusses the FDAP Stack (Flight, DataFusion, Arrow & Parquet), the shift from row-based to columnar data storage, and how technologies like Apache Iceberg are enabling a new era of interoperability across data platforms. The discussion covers why this modular approach saves years of development time while providing better performance and compatibility.RECOMMENDED BOOKSKasun Indrasiri & Danesh Kuruppu • gRPC: Up and Running • https://amzn.to/3sBGBJJTomer Shiran, Jason Hughes & Alex Merced • Apache Iceberg: The Definitive Guide • https://amzn.to/488Z30kWilliam Smith • Arrow Flight Protocols and Practices • https://amzn.to/4o2Q2fdMatthew Topol • In-Memory Analytics with Apache Arrow • https://amzn.to/4oJQ6BMApache Parquet A Complete Guide • https://amzn.to/4i7HVN6BlueskyTwitterInstagramLinkedInFacebookCHANNEL MEMBERSHIP BONUSJoin this channel to get early access to videos & other perks:https://www.youtube.com/channel/UCs_tLP3AiwYKwdUHpltJPuA/joinLooking for a unique learning experience?Attend the next GOTO conference near you! Get your ticket: gotopia.techSUBSCRIBE TO OUR YOUTUBE CHANNEL - new videos posted daily!
SQLite is embedded everywhere - phones, browsers, IoT devices. It's reliable, battle-tested, and feature-rich. But what if you want concurrent writes? Or CDC for streaming changes? Or vector indexes for AI workloads? The SQLite codebase isn't accepting new contributors, and the test suite that makes it so reliable is proprietary. So how do you evolve an embedded database that's effectively frozen?Glauber Costa spent a decade contributing to the Linux kernel at Red Hat, then helped build Scylla, a high-performance rewrite of Cassandra. Now he's applying those lessons to SQLite. After initially forking SQLite (which produced a working business but failed to attract contributors), his team is taking the bolder path: a complete rewrite in Rust called Turso. The project already has features SQLite lacks - vector search, CDC, browser-native async operation - and is using deterministic simulation testing (inspired by TigerBeetle) to match SQLite's legendary reliability without access to its test suite.The conversation covers why rewrites attract contributors where forks don't, how the Linux kernel maintains quality with thousands of contributors, why Pekka's "pet project" jumped from 32 to 64 contributors in a month, and what it takes to build concurrent writes into an embedded database from scratch.--Support Developer Voices on Patreon: https://patreon.com/DeveloperVoicesSupport Developer Voices on YouTube: https://www.youtube.com/@DeveloperVoices/joinTurso: https://turso.tech/Turso GitHub: https://github.com/tursodatabase/tursolibSQL (SQLite fork): https://github.com/tursodatabase/libsqlSQLite: https://www.sqlite.org/Rust: https://rust-lang.org/ScyllaDB (Cassandra rewrite): https://www.scylladb.com/Apache Cassandra: https://cassandra.apache.org/DuckDB (analytical embedded database): https://duckdb.org/MotherDuck (DuckDB cloud): https://motherduck.com/dqlite (Canonical distributed SQLite): https://canonical.com/dqliteTigerBeetle (deterministic simulation testing): https://tigerbeetle.com/Redpanda (Kafka alternative): https://www.redpanda.com/Linux Kernel: https://kernel.org/Datadog: https://www.datadoghq.com/Glauber Costa on X: https://x.com/glcstGlauber Costa on GitHub: https://github.com/glommerKris on Bluesky: https://bsky.app/profile/krisajenkins.bsky.socialKris on Mastodon: http://mastodon.social/@krisajenkinsKris on LinkedIn: https://www.linkedin.com/in/krisjenkins/--0:00 Intro3:16 Ten Years Contributing to the Linux Kernel15:17 From Linux to Startups: OSv and Scylla26:23 Lessons from Scylla: The Power of Ecosystem Compatibility33:00 Why SQLite Needs More37:41 Open Source But Not Open Contribution48:04 Why a Rewrite Attracted Contributors When a Fork Didn't57:22 How Deterministic Simulation Testing Works1:06:17 70% of SQLite in Six Months1:12:12 Features Beyond SQLite: Vector Search, CDC, and Browser Support1:19:15 The Challenge of Adding Concurrent Writes1:25:05 Building a Self-Sustaining Open Source Community1:30:09 Where Does Turso Fit Against DuckDB?1:41:00 Could Turso Compete with Postgres?1:46:21 How Do You Avoid a Toxic Community Culture?1:50:32 Outro
Marimo is redefining what a Python notebook can do—bringing structure, version control, and interactivity together. In this episode, we chat with Akshay Agrawal, co-founder and CEO of Marimo, about how their reactive Python notebook fixes hidden state, keeps outputs in sync, and makes reproducible, reviewable code the norm.Akshay shares Marimo's origin story, how its reactive DAG turns notebooks into clean, Git-friendly tools, and why teams are ditching Jupyter-to-Streamlit pipelines for simpler, reactive workflows. We also dive into performance, data handling with pandas/Polars via Narwhals, and SQL reactivity with DuckDB.Join us in this insightful episode as we talk with Akshay about reproducibility, data workflows, and turning prototypes into shareable apps.For more info on Marimo, reach out to Akshay:Website: https://www.akshayagrawal.com/Github: https://github.com/akshaykaLinkedIn: https://www.linkedin.com/in/akshayka/X: https://x.com/akshaykagrawal______If you found this podcast helpful, please consider following us!Start Here with Pybites: https://pybit.esDeveloper Mindset Newsletter: https://pybit.es/newsletter
Jordan Tigani, CEO and cofounder of MotherDuck, knows what world class infrastructure looks like. He spent years building Google BigQuery before taking those lessons into the startup world. In this episode, he breaks down why building infrastructure products is fundamentally different from typical SaaS and why founders who don't understand that difference are in for a painful surprise.What You'll LearnThere are no shortcuts in infrastructure. You can't just wire together existing open source components and call it a product. Real infrastructure requires contributing meaningfully to the state of the art, and that takes time, money, and deeper technical investment than most founders expect.Starting with startups, not enterprises, is often the smarter play. Early stage infrastructure companies should target other startups first because they're more comfortable with bleeding edge tech, have lower security barriers, and won't force you to spend three engineers building custom auth instead of your actual product.Scaling down is the new scaling up. Jordan saw pressure at SingleStore to make databases smaller and more efficient, not just bigger. That insight led to MotherDuck, which is built on DuckDB—a database that can run in a car, scale to massive cloud instances, and challenge the coordination overhead of legacy distributed systems.Bottoms up engineering cultures win in infrastructure. At BigQuery, engineers close to customer problems could ship fast and independently. Jordan's recreating that at MotherDuck by removing layers between engineers and customers, because creative problem solving requires understanding business constraints, not just technical ones.Convincing people you can scale is half the battle. The best proof is customers who look like your next target and can vouch for you. Next best is real data and benchmarks. If you don't have those yet, lean on implementation support and help prospects test at scale themselves. Early on, sometimes all you have is your word.Timestamped Highlights[01:22] Why infrastructure takes longer to build than typical SaaS products and why there's no shallow way to do it[06:57] The MVP dilemma: finding product market fit when enterprises demand reliability from day one[11:44] Lessons from BigQuery and SingleStore—what to carry over from big tech and what to leave behind[21:21] The gap in the market that led to MotherDuck: why distributed databases don't scale down and why that matters now[26:10] Redefining scale: why 100 users on one giant instance isn't necessarily better than 100 auto scaling individual instances[29:08] The hierarchy of proof: from customer testimonials to benchmarks to trust me, it'll workA Line to Remember“If you really want to build an infrastructure product, you can't just string existing components together. You actually have to contribute meaningfully to improving the state of the art.”Stay ConnectedIf this breakdown of infrastructure startups resonated with you, subscribe so you don't miss future episodes. And if you're building in this space or thinking about it, connect with Jordan on LinkedIn. He's committed to paying forward the help he got as a founder.
0:00 - Motivation for Microsoft to compete with Esri1:20 - Who is Rakesh and data engineering services of SketchMyView6:40 - Deploying a land and planning GIS for the UK government with Microsoft Synapse 22:25 - Is there a similarity between Synapse and Fabric?26:26 - ACID compliance, delta files, lakehouses, bronze, silver, gold layers35:18 - Apache Sedona in Fabric tutorial57:20 - Why is it worth it to use Apache Sedona in Fabric?1:00:44 - SedonaDBApache Sedona is a way for a regular Apache Spark using data analyst to acquire geospatial capabilities. With Sedona, if you know SQL, you know GIS. Rakesh Gupta is Principal Consultant at SketchMyView in London. He tells us about how to set up Apache Sedona in Microsoft Fabric in 2 lines of code. It was a privilege to have his time for this tutorial as he showed how easy it is to get up and running with a powerful, free spatial analysis system that leverages Apache Spark for scalable compute. He also touched in the new SedonaDB, released last month. This is a significant development for the geospatial economy because it is a database created with geospatial data as a first class citizen. This means we have our own database library that is only a pip install away:pip install "apache-sedona[db]"Something to consider as a replacement for DuckDB. More here and here.
Talk Python To Me - Python conversations for passionate developers
Python in 2025 is different. Threads really are about to run in parallel, installs finish before your coffee cools, and containers are the default. In this episode, we count down 38 things to learn this year: free-threaded CPython, uv for packaging, Docker and Compose, Kubernetes with Tilt, DuckDB and Arrow, PyScript at the edge, plus MCP for sane AI workflows. Expect practical wins and migration paths. No buzzword bingo, just what pays off in real apps. Join me along with Peter Wang and Calvin Hendrix-Parker for a fun, fast-moving conversation. Episode sponsors Seer: AI Debugging, Code TALKPYTHON Agntcy Talk Python Courses Links from the show Calvin Hendryx-Parker: github.com/calvinhp Peter on BSky: @wang.social Free-Threaded Wheels: hugovk.github.io Tilt: tilt.dev The Five Demons of Python Packaging That Fuel Our ...: youtube.com Talos Linux: talos.dev Docker: Accelerated Container Application Development: docker.com Scaf - Six Feet Up: sixfeetup.com BeeWare: beeware.org PyScript: pyscript.net Cursor: The best way to code with AI: cursor.com Cline - AI Coding, Open Source and Uncompromised: cline.bot Watch this episode on YouTube: youtube.com Episode #524 deep-dive: talkpython.fm/524 Episode transcripts: talkpython.fm Theme Song: Developer Rap
Mark Raasveldt, co-founder and CTO of DuckDB Labs, shares his journey from academic research at CWI Amsterdam to creating one of the most innovative analytical databases of the last decade. Mark discusses the technical challenges of building DuckDB from scratch, the philosophy behind embedded analytical databases, and why single-node performance still matters in our cloud-first world. He provides insights into open source business models, the evolution of data formats like Parquet, and how DuckDB is democratizing high-performance analytics for developers everywhere.
At PyData Berlin, community members and industry voices highlighted how AI and data tooling are evolving across knowledge graphs, MLOps, small-model fine-tuning, explainability, and developer advocacy.- Igor Kvachenok (Leuphana University / ProKube) combined knowledge graphs with LLMs for structured data extraction in the polymer industry, and noted how MLOps is shifting toward LLM-focused workflows.- Selim Nowicki (Distill Labs) introduced a platform that uses knowledge distillation to fine-tune smaller models efficiently, making model specialization faster and more accessible.- Gülsah Durmaz (Architect & Developer) shared her transition from architecture to coding, creating Python tools for design automation and volunteering with PyData through PyLadies.- Yashasvi Misra (Pure Storage) spoke on explainable AI, stressing accountability and compliance, and shared her perspective as both a data engineer and active Python community organizer.- Mehdi Ouazza (MotherDuck) reflected on developer advocacy through video, workshops, and branding, showing how creative communication boosts adoption of open-source tools like DuckDB.Igor KvachenokMaster's student in Data Science at Leuphana University of Lüneburg, writing a thesis on LLM-enhanced data extraction for the polymer industry. Builds RDF knowledge graphs from semi-structured documents and works at ProKube on MLOps platforms powered by Kubeflow and Kubernetes.Connect: https://www.linkedin.com/in/igor-kvachenok/Selim NowickiFounder of Distill Labs, a startup making small-model fine-tuning simple and fast with knowledge distillation. Previously led data teams at Berlin startups like Delivery Hero, Trade Republic, and Tier Mobility. Sees parallels between today's ML tooling and dbt's impact on analytics.Connect: https://www.linkedin.com/in/selim-nowicki/Gülsah DurmazArchitect turned developer, creating Python-based tools for architectural design automation with Rhino and Grasshopper. Active in PyLadies and a volunteer at PyData Berlin, she values the community for networking and learning, and aims to bring ML into architecture workflows.Connect: https://www.linkedin.com/in/gulsah-durmaz/Yashasvi (Yashi) MisraData Engineer at Pure Storage, community organizer with PyLadies India, PyCon India, and Women Techmakers. Advocates for inclusive spaces in tech and speaks on explainable AI, bridging her day-to-day in data engineering with her passion for ethical ML.Connect: https://www.linkedin.com/in/misrayashasvi/Mehdi OuazzaDeveloper Advocate at MotherDuck, formerly a data engineer, now focused on building community and education around DuckDB. Runs popular YouTube channels ("mehdio DataTV" and "MotherDuck") and delivered a hands-on workshop at PyData Berlin. Blends technical clarity with creative storytelling.Connect: https://www.linkedin.com/in/mehd-io/
The DuckLake Lakehouse Format // MLOps Podcast #339 with Hannes Mühleisen, Co-founder and CEO of DuckDB Labs.Join the Community: https://go.mlops.community/YTJoinInGet the newsletter: https://go.mlops.community/YTNewsletter// AbstractManaging data on Object Stores has been a painful affair. Users had to choose between data swamp chaos or a maze of metadata files with catalog servers on top. DuckLake is a new paradigm for managing data on object stores: First, it uses classical SQL data management systems to manage metadata. Second, actual data is stored in Parquet files on pretty arbitrary storage. Third, processing queries is done client-side, or anywhere really. DuckDB is the first system to integrate with DuckLake using an extension with the same name. Conceptually, DuckLake enables central control over truth while decentralizing compute and storage entirely. DuckLake turns data warehouse architecture upside down by departing from the integrated metadata/compute layer towards a fully disconnected operation with only centralized metadata. For the first time, DuckLake allows a “multi-player” experience with DuckDB, where computation stays fully local, but transactional control is centralized.// BioHannes Mühleisen
In this episode of Stories from the Hackery, we talk with Nashville tech leader and hiring manager Jason Turan about one of tech's most in-demand fields: data engineering. Jason, a long-time friend of NSS, was one of the first people to tell us that Nashville needed more data engineers. He shares his perspective on what a data engineer does, describing the role as the "connective tissue between data producers and data consumers". Listen in to hear us discuss: - Why data engineers are essential for flipping the 80/20 rule, allowing data scientists and analysts to spend less time cleaning data and more time finding insights. - How the rise of generative AI has acted as an "accelerant," increasing the need for high-quality data and the professionals who can provide it. - Actionable advice for getting started in the field, including the importance of focusing on a "T-shaped skillset" with SQL at its core. - Why Jason's number one piece of advice is to be curious, experiment, and "go out and do the thing". 01:20 Meet Jason Turan: His Tech Origin Story 03:04 Jason's History with NSS and Hiring Grads 07:28 Defining Data Engineering: The "Connective Tissue" of Tech 11:15 Why Nashville is a Hub for Data Engineers 13:56 Healthcare's Impact on Nashville's Data Jobs 20:35 How GenAI Accelerates the Need for Data Engineers 31:33 Getting Started: Lower Barriers to Entry 39:03 A Top Use Case for AI: Understanding Your Codebase 52:21 Misconceptions & the "T-Shaped Skillset" 55:29 The Value of Hands-On Learning: "Go Do the Thing" 58:52 Lightning Round: Favorite Tech Tools 01:00:32 Lightning Round: Top Reads & Resources Links Metabase: https://www.metabase.com/ DuckDB: https://duckdb.org/ MotherDuck: https://motherduck.com/ Ralph Kimball: The Data Warehouse Toolkit: https://www.amazon.com/gp/product/1118530802 Bill Inmon: Building the Data Warehouse: https://www.amazon.com/Building-Data-Warehouse-W-Inmon/dp/0764599445 Edward Tufte: The Visual Display of Quantitative Information: https://www.amazon.com/Visual-Display-Quantitative-Information/dp/0961392142 Brendan Keeler: The Health API Guy: https://healthapiguy.substack.com/ TLDR Newsletter: https://tldr.tech/ Nashville Technology Council (NTC): https://technologycouncil.com/
In this episode, we talk with Orell about his journey from electrical engineering to freelancing in data engineering. Exploring lessons from startup life, working with messy industrial data, the realities of freelancing, and how to stay up to date with new tools. Topics covered: Why Orel left a PhD and a simulation‑focused start‑up after Covid hitWhat he learned trying (and failing) to commercialise medical‑imaging simulationsThe first freelance project and the long, quiet months that followedHow he now finds clients, keeps projects small and delivers value quicklyTypical work he does for industrial companies: parsing messy machine logs, building simple pipelines, adding structure laterFavorite everyday tools (Python, DuckDB, a bit of C++) and the habit of blocking time for learningAdvice for anyone thinking about freelancing: cash runway, networking, and focusing on problems rather than “perfect” tech choicesA practical conversation for listeners who are curious about moving from research or permanent roles into freelance data engineering.
Topics covered in this episode: * Distributed sqlite follow up: Turso and Litestream* * PEP 792 – Project status markers in the simple index* Run coverage on tests docker2exe: Convert a Docker image to an executable Extras Joke Watch on YouTube About the show Sponsored by Digital Ocean: pythonbytes.fm/digitalocean-gen-ai Use code DO4BYTES and get $200 in free credit Connect with the hosts Michael: @mkennedy@fosstodon.org / @mkennedy.codes (bsky) Brian: @brianokken@fosstodon.org / @brianokken.bsky.social Show: @pythonbytes@fosstodon.org / @pythonbytes.fm (bsky) Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Monday at 10am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Michael #1: Distributed sqlite follow up: Turso and Litestream Michael Booth: Turso marries the familiarity and simplicity of SQLite with modern, scalable, and distributed features. Seems to me that Turso is to SQLite what MotherDuck is to DuckDB. Mike Fiedler Continue to use the SQLite you love and care about (even the one inside Python runtime) and launch a daemon that watches the db for changes and replicates changes to an S3-type object store. Deeper dive: Litestream: Revamped Brian #2: PEP 792 – Project status markers in the simple index Currently 3 status markers for packages Trove Classifier status Indices can be yanked PyPI projects - admins can quarantine a project, owners can archive a project Proposal is to have something that can have only one state active archived quarantined deprecated This has been Approved, but not Implemented yet. Brian #3: Run coverage on tests Hugo van Kemenade And apparently, run Ruff with at least F811 turned on Helps with copy/paste/modify mistakes, but also subtler bugs like consumed generators being reused. Michael #4: docker2exe: Convert a Docker image to an executable This tool can be used to convert a Docker image to an executable that you can send to your friends. Build with a simple command: $ docker2exe --name alpine --image alpine:3.9 Requires docker on the client device Probably doesn't map volumes/ports/etc, though could potentially be exposed in the dockerfile. Extras Brian: Back catalog of Test & Code is now on YouTube under @TestAndCodePodcast So far 106 of 234 episodes are up. The rest are going up according to daily limits. Ordering is rather chaotic, according to upload time, not release ordering. There will be a new episode this week pytest-django with Adam Johnson Joke: If programmers were doctors
The three of us talk with Christoph Windheuser about the styles in data architecture: data mesh, data lake (house) and data warehouse and how to make a decision. In between Christoph explains data quality, data lineage, and data catalog - cornerstones of any modern approach. We end with emerging trends, DuckDB and data governance.
This week on The Data Stack Show, John Wessel and Matt Kelliher-Gibson dive into the recent Duck Lake announcement, exploring the evolving landscape of data analytics technologies. They discuss DuckDB's role as a lightweight, local analytics database and its potential as a caching layer for open table formats like Iceberg. The conversation also highlights the current state of data storage standards, focusing on agreements around Parquet and Iceberg, while noting the ongoing complexity in catalog management. Key takeaways include the importance of local compute solutions, the early stage of open table formats, and the potential for simplified data infrastructure that can provide faster, more cost-effective analytics workflows. The episode underscores the ongoing innovation in data technologies and the need for more streamlined, flexible data management solutions. Don't miss it!Highlights from this week's conversation include:Discussion on Duck Lake Announcement (1:41)Compatibility with Apache Iceberg (4:05)Use Cases for DuckDB (6:23)Concerns About Data Management (10:01)Introduction to Data Formats (11:40)Catalog Space Challenges (13:13)Metadata Orchestration (14:54)Simplicity in Data Management (15:25)SQL Demo Discussion (17:26)Wrap-Up and Final Thoughts (18:44)The Data Stack Show is a weekly podcast powered by RudderStack, customer data infrastructure that enables you to deliver real-time customer event data everywhere it's needed to power smarter decisions and better customer experiences. Each week, we'll talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data.RudderStack helps businesses make the most out of their customer data while ensuring data privacy and security. To learn more about RudderStack visit rudderstack.com.
Imagine writing SQL and getting instant results as you type? Yes, this is reality now. It's amazing!DuckDB/MotherDuck's Instant SQL made a big splash at last month's Data Council. Hamilton Ulmer gives a demo of Instant SQL at the Practical Data Community.----------------------------Instant SQL: https://motherduck.com/blog/introducing-instant-sql/Practical Data Community Discord: https://discord.gg/gNfw5AKWSK
Are you looking for a fast database that can handle large datasets in Python? What's the difference between a Python expression and a statement? Christopher Trudeau is back on the show this week, bringing another batch of PyCoder's Weekly articles and projects.
In this episode of The Data Engineering Show, the bros welcome the CEO DuckDB Labs and co-creator DuckDB, Hannes Mühleisen. They delve into the groundbreaking journey of DuckDB, an analytical database that processes billions of queries every month. Learn why DuckDB prioritizes broad compatibility over specialized optimizations, how its extension model works and the emerging solutions for database technology in the age of AI.
In this podcast episode, we talked with Adrian Brudaru about the past, present and future of data engineering.About the speaker:Adrian Brudaru studied economics in Romania but soon got bored with how creative the industry was, and chose to go instead for the more factual side. He ended up in Berlin at the age of 25 and started a role as a business analyst. At the age of 30, he had enough of startups and decided to join a corporation, but quickly found out that it did not provide the challenge he wanted.As going back to startups was not a desirable option either, he decided to postpone his decision by taking freelance work and has never looked back since. Five years later, he co-founded a company in the data space to try new things. This company is also looking to release open source tools to help democratize data engineering.0:00 Introduction to DataTalks.Club1:05 Discussing trends in data engineering with Adrian2:03 Adrian's background and journey into data engineering5:04 Growth and updates on Adrian's company, DLT Hub9:05 Challenges and specialization in data engineering today13:00 Opportunities for data engineers entering the field15:00 The "Modern Data Stack" and its evolution17:25 Emerging trends: AI integration and Iceberg technology27:40 DuckDB and the emergence of portable, cost-effective data stacks32:14 The rise and impact of dbt in data engineering34:08 Alternatives to dbt: SQLMesh and others35:25 Workflow orchestration tools: Airflow, Dagster, Prefect, and GitHub Actions37:20 Audience questions: Career focus in data roles and AI engineering overlaps39:00 The role of semantics in data and AI workflows41:11 Focusing on learning concepts over tools when entering the field 45:15 Transitioning from backend to data engineering: challenges and opportunities 47:48 Current state of the data engineering job market in Europe and beyond 49:05 Introduction to Apache Iceberg, Delta, and Hudi file formats 50:40 Suitability of these formats for batch and streaming workloads 52:29 Tools for streaming: Kafka, SQS, and related trends 58:07 Building AI agents and enabling intelligent data applications 59:09Closing discussion on the place of tools like DBT in the ecosystem
A major milestone for leveraging LLMs in R just landed with the new ellmer package, along with a terrific showcase of retrieval-augmented generation combining ellmer and DuckDB. Plus an inspiring roundup of the recent Closeread contest winners.Episode LinksThis week's curator: Sam Parmar - @parmsam@fosstodon.org (Mastodon) & @parmsam_ (X/Twitter)Announcing ellmer: A package for interacting with Large Language Models in RRapid RAG Prototyping: Building a Retrieval Augmented Generation Prototype with ellmer and DuckDBWinners of the Closeread Prize – Data-Driven Scrollytelling with QuartoEntire issue available at rweekly.org/2025-W10Supplement ResourcesCoder Radio episode 608 - R with Eric Nantz https://coder.show/608nhyris - The minimal framework for transform R shiny application into standaloneSupporting the showUse the contact page at https://serve.podhome.fm/custompage/r-weekly-highlights/contact to send us your feedbackR-Weekly Highlights on the Podcastindex.org - You can send a boost into the show directly in the Podcast Index. First, top-up with Alby, and then head over to the R-Weekly Highlights podcast entry on the index.A new way to think about value: https://value4value.infoGet in touch with us on social mediaEric Nantz: @rpodcast@podcastindex.social (Mastodon), @rpodcast.bsky.social (BlueSky) and @theRcast (X/Twitter)Mike Thomas: @mike_thomas@fosstodon.org (Mastodon), @mike-thomas.bsky.social (BlueSky), and @mike_ketchbrook (X/Twitter) Music credits powered by OCRemixWatermelon Flava - Breath of Fire III - Joshua Morse, posu yan - https://ocremix.org/remix/OCR01411Stomp the Summer Sky - Secret of Mana - Ziwtra - https://ocremix.org/remix/OCR00859
Michael and Nikolay are joined by Joe Sciarrino and Jelte Fennema-Nio to discuss pg_duckdb — what it is, how it started, what early users are using it for, and what they're working on next. Here are some links to things they mentioned:Joe Sciarrino https://postgres.fm/people/joe-sciarrinoJelte Fennema-Nio https://postgres.fm/people/jelte-fennema-niopg_duckdb https://github.com/duckdb/pg_duckdbHydra https://www.hydra.soMotherDuck https://motherduck.comThe problems and benefits of an elephant with a beak (lightning talk by Jelte) https://www.youtube.com/watch?v=ogvbKE4fw9A&list=PLF36ND7b_WU4QL6bA28NrzBOevqUYiPYq&t=1073spg_duckdb announcement post (by Jordan and Brett from MotherDuck) https://motherduck.com/blog/pg_duckdb-postgresql-extension-for-duckdb-motherduckpg_duckdb 0.2 release https://github.com/duckdb/pg_duckdb/releases/tag/v0.2.0~~~What did you like or not like? What should we discuss next time? Let us know via a YouTube comment, on social media, or by commenting on our Google doc!~~~Postgres FM is produced by:Michael Christofides, founder of pgMustardNikolay Samokhvalov, founder of Postgres.aiWith special thanks to:Jessie Draws for the elephant artwork
Talk Python To Me - Python conversations for passionate developers
Join me for an insightful conversation with Alex Monahan, who works on documentation, tutorials, and training at DuckDB Labs. We explore why DuckDB is gaining momentum among Python and data enthusiasts, from its in-process database design to its blazingly fast, columnar architecture. We also dive into indexing strategies, concurrency considerations, and the fascinating way MotherDuck (the cloud companion to DuckDB) handles large-scale data seamlessly. Don't miss this chance to learn how a single pip install could totally transform your Python data workflow! Episode sponsors Sentry Error Monitoring, Code TALKPYTHON Data Citizens Podcast Talk Python Courses Links from the show Alex on Mastodon: @__Alex__ DuckDB: duckdb.org MotherDuck: motherduck.com SQLite: sqlite.org Moka-Py: github.com PostgreSQL: www.postgresql.org MySQL: www.mysql.com Redis: redis.io Apache Parquet: parquet.apache.org Apache Arrow: arrow.apache.org Pandas: pandas.pydata.org Polars: pola.rs Pyodide: pyodide.org DB-API (PEP 249): peps.python.org/pep-0249 Flask: flask.palletsprojects.com Gunicorn: gunicorn.org MinIO: min.io Amazon S3: aws.amazon.com/s3 Azure Blob Storage: azure.microsoft.com/products/storage Google Cloud Storage: cloud.google.com/storage DigitalOcean: www.digitalocean.com Linode: www.linode.com Hetzner: www.hetzner.com BigQuery: cloud.google.com/bigquery DBT (Data Build Tool): docs.getdbt.com Mode: mode.com Hex: hex.tech Python: www.python.org Node.js: nodejs.org Rust: www.rust-lang.org Go: go.dev .NET: dotnet.microsoft.com Watch this episode on YouTube: youtube.com Episode transcripts: talkpython.fm --- Stay in touch with us --- Subscribe to Talk Python on YouTube: youtube.com Talk Python on Bluesky: @talkpython.fm at bsky.app Talk Python on Mastodon: talkpython Michael on Bluesky: @mkennedy.codes at bsky.app Michael on Mastodon: mkennedy
Hannes Muhleisen is the creator of DuckDB and CEO of DuckDB Labs. We finally got a chance to meet in person at the Forward Data Conference in Paris. We hit it off immediately, and at times, I felt like I was talking with my long lost brother. Hannes is a very cool guy! While at the conference, we recorded a chat about all things DuckDB, the challenges of data lakehouses and open table formats, local-first tech, and much more.
We are on the other side of "big data" hype, but what is the future of analytics and how does AI fit in? Till and Adithya from MotherDuck join us to discuss why DuckDB is taking the analytics and AI world by storm. We dive into what makes DuckDB, a free, in-process SQL OLAP database management system, unique including its ability to execute lighting fast analytics queries against a variety of data sources, even on your laptop! Along the way we dig into the intersections with AI, such as text-to-sql, vector search, and AI-driven SQL query correction.
A founding engineer on Google BigQuery and now at the helm of MotherDuck, Jordan Tigani challenges the decade-long dominance of Big Data and introduces a compelling alternative that could change how companies handle data. Jordan discusses why Big Data technologies are an overkill for most companies, how MotherDuck and DuckDB offer fast analytical queries, and lessons learned as a technical founder building his first startup. Watch the episode with Tomasz Tunguz: https://youtu.be/gU6dGmZzmvI Website - https://motherduck.com Twitter - https://x.com/motherduck Jordan Tigani LinkedIn - https://www.linkedin.com/in/jordantigani Twitter - https://x.com/jrdntgn FIRSTMARK Website - https://firstmark.com Twitter - https://twitter.com/FirstMarkCap Matt Turck (Managing Director) LinkedIn - https://www.linkedin.com/in/turck/ Twitter - https://twitter.com/mattturck (00:00) Intro (00:56) What is the Small Data? (06:56) Marketing strategy of MotherDuck (08:39) Processing Small Data with Big Data stack (15:30) DuckDB (17:21) Creation of DuckDB (18:48) Founding story of MotherDuck (24:08) MotherDuck's community (25:25) MotherDuck of today ($100M raised) (33:15) Why MotherDuck and DuckDB are so fast? (39:08) The limitations and the future of MotherDuck's platform (39:49) Small Models (42:37) Small Data and the Modern Data Stack (46:47) Making things simpler with a shift from Big Data to Small Data (50:04) Jordan Tigani's entrepreneurial journey (58:31) Outro
В этом выпуске мы делимся еженедельными открытиями, обсуждаем VPN в России, сравниваем Swift и Rust, говорим о DirectX 9, Windows10, DuckDB 1.1.0 и ретрогейминге. [00:03:22] Чемы мы научились на этой неделе The first professional hosting of cloud VPS/VDS servers — VDSina Open Data Protocol — Wikipedia Сварочный инвертор за 5$ своими руками! https://www.amazon.co.uk/dp/B0C9WWCQ82/ref=emc_bcc_2_i?th=1 [00:03:39] VPN который… Читать далее →
DuckDB is an open-source column-oriented relational database that was first released in 2019. It's designed to provide high performance on complex queries against large databases, and focuses on online analytical processing workloads. Hannes Mühleisen is the Co-Creator of DuckBD, and is the CEO and Co-Founder of DuckDB Labs. He joins the show to talk about The post DuckDB with Hannes Mühleisen appeared first on Software Engineering Daily.
Topics covered in this episode: PSF Elections coming up Cloud engineer gets 2 years for wiping ex-employer's code repos Python: Import by string with pkgutil.resolve_name() DuckDB goes 1.0 Extras Joke Watch on YouTube About the show Sponsored by ScoutAPM: pythonbytes.fm/scout Connect with the hosts Michael: @mkennedy@fosstodon.org Brian: @brianokken@fosstodon.org Show: @pythonbytes@fosstodon.org Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesdays at 10am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Brian #1: PSF Elections coming up This is elections for the PSF Board and for 3 bylaw changes. To vote in the PSF election, you need to be a Supporting, Managing, Contributing, or Fellow member of the PSF, … And affirm your voting status by June 25. See Affirm your PSF Membership Voting Status for more details. Timeline Board Nominations open: Tuesday, June 11th, 2:00 pm UTC Board Nominations close: Tuesday, June 25th, 2:00 pm UTC Voter application cut-off date: Tuesday, June 25th, 2:00 pm UTC same date is also for voter affirmation. Announce candidates: Thursday, June 27th Voting start date: Tuesday, July 2nd, 2:00 pm UTC Voting end date: Tuesday, July 16th, 2:00 pm UTC See also Thinking about running for the Python Software Foundation Board of Directors? Let's talk! There's still one upcoming office hours session on June 18th, 12 PM UTC And For your consideration: Proposed bylaws changes to improve our membership experience 3 proposed bylaws changes Michael #2: Cloud engineer gets 2 years for wiping ex-employer's code repos Miklos Daniel Brody, a cloud engineer, was sentenced to two years in prison and a restitution of $529,000 for wiping the code repositories of his former employer in retaliation for being fired. The court documents state that Brody's employment was terminated after he violated company policies by connecting a USB drive. Brian #3: Python: Import by string with pkgutil.resolve_name() Adam Johnson You can use pkgutil.resolve_name("[HTML_REMOVED]:[HTML_REMOVED]")to import classes, functions or modules using strings. You can also use importlib.import_module("[HTML_REMOVED]") Both of these techniques are so that you have an object imported, but the end thing isn't imported into the local namespace. Michael #4: DuckDB goes 1.0 via Alex Monahan The cloud hosted product @MotherDuck also opened up General Availability Codenamed "Snow Duck" The core theme of the 1.0.0 release is stability. Extras Brian: Sending us topics. Please send before Tuesday. But any time is welcome. NumPy 2.0 htmx 2.0.0 Michael: Get 6 months of PyCharm Pro for free. Just take a course (even a free one) at Talk Python Training. Then visit your account page > details tab and have fun. Coming soon at Talk Python: Shiny for Python Joke: .gitignore thoughts won't let me sleep
How do you debug your EF queries? Carl and Richard talk to Giorgi Dalakishvili about his open-source Visual Studio extension, EFCore Visualizer. Giorgi talks about bringing together the EF rendering of the query with the database query plan to ensure you retrieve data from your database as efficiently as possible. The conversation ranges over a number of tools Giorgi has built over the years, including EF Framework Exceptions, DuckDB.NET, and more!
Talk Python To Me - Python conversations for passionate developers
Do you have data that you pull from external sources or is generated and appears at your digital doorstep? I bet that data needs processed, filtered, transformed, distributed, and much more. One of the biggest tools to create these data pipelines with Python is Dagster. And we are fortunate to have Pedram Navid on the show this episode. Pedram is the Head of Data Engineering and DevRel at Dagster Labs. And we're talking data pipelines this week at Talk Python. Episode sponsors Talk Python Courses Posit Links from the show Rock Solid Python with Types Course: training.talkpython.fm Pedram on Twitter: twitter.com Pedram on LinkedIn: linkedin.com Ship data pipelines with extraordinary velocity: dagster.io dagster-open-platform: github.com The Dagster Master Plan: dagster.io data load tool (dlt): dlthub.com DataFrames for the new era: pola.rs Apache Arrow: arrow.apache.org DuckDB is a fast in-process analytical database: duckdb.org Ship trusted data products faster: www.getdbt.com Watch this episode on YouTube: youtube.com Episode transcripts: talkpython.fm --- Stay in touch with us --- Subscribe to us on YouTube: youtube.com Follow Talk Python on Mastodon: talkpython Follow Michael on Mastodon: mkennedy