Language for management and use of relational databases
POPULARITY
Categories
Hey friends! We've been on a bit of a Tales of Pentest Pwnage bender lately, so let's keep it rolling with Part 90. (And Mom, relax — this is not one pentest story chopped into 90 parts.) Today is less of an A-to-Z story and more a pile of tips and tricks pulled from a recent string of SCCM-flavored internals — plus a tangent about a video game and a little robot Claude and I built to get my life back. Multi-tier SCCM is having a moment — I've never administered SCCM a day in my life, but 2026 keeps dropping me into these split-role environments. Here's where I go to figure out my attack surface when all I've got is a low-priv cred. SMB signing on? Cool, I'll go around it — relaying from one SCCM box to the one with SQL on it, and the easy-button tool that vacuums the good stuff out of the database once you're there. (NAA creds, clear-text local admin passwords, install scripts with creds baked in…yes please.) Then relay the other direction — when the loot came back stale, a nudge from a Slack channel had me pointing the relay backwards, and I giggled like a little schoolboy at what popped out. Adding yourself to the local admin group: still weirdly undetected — an old episode of ours reminded me of an Impacket tool I don't see written up on many pentest blogs, and it slipped right by. My favorite cheat code hit a snag — the evil-scheduled-task-under-a-logged-in-DA trick kept coughing up permission errors, so I had to get creative. Plus the defensive recs I'm still trying to sharpen up — if you've got a better way to close this loophole, I'm all ears! Tangent: Halloween (the video game) and the lobby watcher — a million players and I'm still staring into the lobby abyss for 10 minutes at a time. So Claude and I built something that watches the screen and texts me when it's time to sprint back to the keyboard. It's not cheating. It's not! (The Texas Chainsaw crowd disagreed, loudly.)
"Czy robiliśmy review lat temu trzy? Nie, nie robiliśmy, tylko wszyscy okłamywali się, że robią review porządnie i klikali: tak, tak, tak może iść."
This week we talk about AI agents, cyberattacks, and insurance claims.We also discuss OpenAI, Hugging Face, and policy language.Recommended Book: The Stars My Destination by Alfred BesterTranscriptTwo broad categories of cyberattack have become especially visible this year, and only one of them requires a human attacker in the loop to choose the target.In March, hackers linked by the US government to Iranian intelligence broke into the medical-device manufacturer Stryker and remotely wiped tens of thousands of employee devices. The attack disrupted the company for days, affected its first-quarter earnings, and represented a shift from somewhat more subtle espionage toward more overt and deliberate destruction.Elsewhere, the market-research company Klue sat at the center of a breach affecting close to 200 customers. Attackers used an old credential to gain access to keys for customers' cloud services.These incidents had people with recognizable motives behind them, and that sort of hack has become increasingly rare over the past decade or so: Black Kite counted 7,551 publicly disclosed ransomware victims over its latest twelve-month reporting period, alone, about 25 percent more than the year before.Then, over the course of about five weeks this summer, OpenAI, Anthropic, and Meta each disclosed incidents in which their AI models gained access to real computer systems they were not supposed to touch.The incidents were not identical, and describing all of them as escapes wouldn't capture the important nuance of the situation.Anthropic and Meta said their models reached the internet because a third-party evaluation environment had been misconfigured. The models had been told they were operating inside a simulation, and their internal dialogues indicate they interpreted the real systems they encountered when they left that environment as simulated components of that exercise.OpenAI's incident, however, was a bit different.Its models were being tested using ExploitGym, a benchmark intended to measure their ability to find and exploit software vulnerabilities. Internet access was supposed to be blocked, but the models discovered and exploited vulnerabilities in a package-registry service inside OpenAI's research infrastructure, found a route to the open internet, and, as I related in a previous episode, eventually compromised the production systems of the AI platform Hugging Face.Since that episode, OpenAI has release a more complete postmortem, which they published in August, and which says the activity was primarily driven by an internal research model comparable in scale to GPT-5.6 Sol, though Sol participated in parts of the incident.The agents created an improvised message board inside the package manager, exchanged discoveries, divided up work, and sometimes described themselves as a swarm.At Hugging Face they executed code on dozens of servers, obtained broad access to several clusters, and exposed credentials associated with databases, cloud services, code repositories, and internal communications.Hugging Face reconstructed around 17,600 actions over a campaign lasting about four and a half days, including roughly two and a half days inside its own infrastructure.Notably, these AI agents were not trying to steal money or damage Hugging Face. They inferred that the company might possess the solutions to the benchmark, to the problem they were tasked with solving, and they went looking for them. They were, in effect, cheating on a test, looking for solutions to the problems rather than taking the test, as intended.That doesn't mean the models had motives in the human sense. People created the evaluation, reduced the normal safeguards, and assigned an offensive task. But no operator selected Hugging Face or directed the individual steps; that was a path the AI decided to take all by itself.Anthropic's review subsequently found three incidents among more than 141,000 evaluation runs for its own models. In one, a model unable to reach its fictional target scanned around 9,000 real ones, compromised an internet-facing application using an exposed password and SQL injection, then stopped after recognizing that the host was real.Another model published a malicious package to the real Python Package Index because it believed the registry was simulated. Fifteen outside systems downloaded and ran that malicious package before it was removed.Meta then disclosed that one of its models had reached the internet through a misconfiguration at the same evaluation vendor and exploited a vulnerability at an unnamed third party.No significant financial damages have been publicly reported from these events, by attacker or victim. But if there had been damages, who would have paid for them?What I'd like to talk about today is how autonomous AI systems complicate cyber insurance, how insurers have handled equally unfamiliar risks in the past, and why insurance contracts may soon become one of the more important forms of AI governance.—A typical cyber-insurance policy covers a broad portfolio of costs.These can include ransom payments, forensic investigations, legal expenses, restoring systems and data, notifying customers, and compensating victims and possibly a victims' customers for the revenue lost while a company's operations are interrupted.Business interruption is often one of the largest portions of a claim, and policies can respond to malicious attacks as well as non-malicious failures.This market grew by more than 30% a year between 2017 and 2022, as ransomware, a type of attack that became a lot more common during that period, in part because of increased automation and a franchising model that became really popular and increased the reach of the most powerful ransomware tools, almost broke this industry.In 2021, attacks on Colonial Pipeline, the insurer CNA, and meat processor JBS produced multimillion-dollar ransom payments and costly disruptions. Insurance prices surged, sometimes by more than 100%, while some companies found they could not obtain coverage because insurers just couldn't make the numbers work for them.Insurers responded to this more complex hacking environment by raising prices, but they also made coverage conditional on specific defenses. Companies increasingly had to demonstrate that they used multifactor authentication, endpoint monitoring, restricted administrator access, and backups that attackers could not alter, as a baseline.Loss ratios then fell, more insurance capital entered the market, and prices eventually came down again, stabilizing after that frantic and uncertain period.According to Marsh, global cyber-insurance rates fell 4% in the second quarter of 2026, the twelfth consecutive quarterly decline. Primary pricing is now about 42% below its 2022 peak.The market is not necessarily becoming safer, though. US cyber premiums reached about $7.5 billion in 2025, while the share of premiums consumed by claims rose to 53%—the first time it ticked above 50% since the pandemic-era ransomware surge.Globally, Munich Re estimates the market was worth nearly $15 billion last year and could approach $28 billion by 2030.During this period, insurance applications have also become a consequential part of a company's security system.In one particularly clear example, Travelers rescinded a million-dollar policy after a ransomware claim revealed that the customer's multifactor authentication protected only its firewall, despite application answers saying the control was used much more broadly.Companies that don't live up to cyber insurance expectations can thus be left in the lurch, so in a very real way, insurers have helped make multifactor authentication a standard business practice by attaching a price to its absence. This industry could move faster than regulators because they didn't have to ban insecure behavior and pass legislation to make that happen; they just had to decline to insure anyone who didn't live up to their basic security standards, which left those who failed to implement such precautions without insurance, should they be targeted by hackers.That same mechanism is now being aimed at AI agents, but the big initial problem everyone is facing is definitional.Most cyber policies are written around some identifiable security event: an outside attacker breaks in, an employee steals information, a credential is used without authorization, or malicious software takes a server offline.What if, though, a company gives an AI agent access to its network so that the agent can find and repair security vulnerabilities?And then maybe the agent discovers a vulnerability, exploits it, moves laterally into systems it was not expected to touch, and exposes sensitive data. There is a cyber loss, but there may be no conventional attacker and no stolen credential. The software was invited in and may have used permissions it was explicitly given. This is very different from a human-led hack, but it still has the potential to cause a lot of monetary damage.Insurers including MSIG, QBE, and Beazley are reviewing how their policy language applies to these scenarios and who bears responsibility when an agent's autonomous actions cause damage.For now, most of them are clarifying the parameters of their coverage rather than excluding AI events entirely.QBE's global head of cyber described AI as “a risk amplifier, not a fundamentally new cyber risk.” In other words, if an AI system causes something that looks like an ordinary covered breach, the involvement of AI probably won't put it in a different category; it'll still be covered.The trickier cases involve an agent that works as designed but makes an expensive decision, or a systemic event in which a model or AI platform produces losses at many companies simultaneously.The first type might be treated as professional liability, or errors and omissions, rather than a cyber incident. The second could, in theory at least, end up being too large for insurers to cover without strict limits in place.Specialized products are already emerging. Armilla AI, Munich Re, and AXA XL sell coverage for risks including model underperformance, hallucinations, and intellectual-property claims. Whether these products remain separate or are eventually folded into broad cyber policies will depend in part on what sorts of claims insurers actually receive, and the scale of those claims.Right now, they have very little historical data with which to calculate the price. Insurance is fundamentally a system for using past experience to account for future issues, and autonomous AI losses have almost no past; they're a very new type of problem.That said, the insurance industry has encountered ambiguity before.For years, insurers worried about silent cyber: losses caused by digital events that appeared inside property, liability, and other policies that had never explicitly contemplated them. Lloyd's gradually required policies to state whether cyber risks were covered or excluded.There was also the question of attribution. In 2017, the NotPetya malware spread from Ukraine through corporate networks around the world. The US and several allies attributed it to the Russian military, but many victims were ordinary companies with no meaningful role in geopolitics.Drugmaker Merck claimed about $1.4 billion in damages under its property policies. Insurers disputed roughly half of that amount using exclusions for hostile or warlike acts; language whose ancestry predates computers by more than a century.New Jersey courts found that the exclusion required something closer to conventional military action, not a cyber bug gone haywire, and the case settled in January of 2024, just before the state Supreme Court was scheduled to hear it. Mondelez settled a parallel dispute with Zurich over a claim exceeding $100 million around the same time.Lloyd's subsequently required standalone cyber policies to address catastrophic state-backed attacks explicitly, including a method for resolving attribution. Agentic AI reopens both the silent-coverage and attribution problems simultaneously.If a model causes a loss, responsibility might plausibly be assigned to the model developer, the company that deployed it, the vendor that built its evaluation environment, or the organization whose excessive permissions allowed the damage to spread.In two of this summer's three sets of incidents, a third-party evaluator's misconfiguration helped create the path to the internet. In OpenAI's case, the path involved vulnerabilities in OpenAI's own infrastructure and then weaknesses at several outside services.The most important insurance risk, though, may ultimately be technological and infrastructural aggregation.The 2024 CrowdStrike outage demonstrated that a single faulty software update could interrupt airlines, banks, hospitals, and other organizations around the world without any malicious attacker.Consider a future in which thousands of companies give access to agents built on a small number of frontier models. A flaw or unwanted behavior in one widely used model could cause problems for a large portion of an insurer's entire customer base, all at once.And this risk is arriving in the midst of an unusually competitive insurance market, after twelve quarters of declining rates and as loss ratios are beginning to rise. If insurers decide they cannot price the exposure, they will probably respond through some combination of higher prices, lower limits, stricter conditions, and exclusions.All that in mind, the first thing to be watching in the coming months is policy language during the January 2027 renewal season.The current posture, if you recall, is to clarify rather than exclude, but language addressing systemic AI events or dependence on a single model provider is already being discussed. A significant loss could change the market's posture quickly, making it more limited and expensive.The second thing to watch for is the first big, disputed claim.Industry interviews can describe what insurers expect to cover, but their operational position will be established when an AI agent causes an eight-figure loss and a carrier must either pay or explain why it won't.The NotPetya disputes took years to resolve, and the first autonomous-agent case could similarly define policy language well before it produces a final court ruling. That'll be a moment that maybe defines the next ten years of cyber insurance standards, if not longer.The third thing to watch for is changes to insurance questionnaires.Underwriters could begin asking whether agent credentials are narrowly scoped, whether actions are comprehensively logged, whether consequential decisions require human approval, whether agents have kill switches, and whether claimed containment has been verified rather than merely documented.If these controls affect the price and availability of insurance, they could become industry standards faster than legislation makes them mandatory, just like that previous round of cyber insurance baselines that became common because, lacking them, customers could no longer get cyber insurance at any price.And finally, there's also a government process developing alongside the private one.An executive order signed in June established a voluntary framework under which developers can provide the federal government with access to certain frontier models for up to 30 days before release. The process uses classified benchmarks to evaluate advanced cyber capabilities, and representatives from major AI companies discussed the framework at the White House in August of 2026.If insurers eventually require evidence that a model or company participated in this sort of evaluation, a voluntary government program could evolve into a practical requirement without ever becoming an actual legal mandate.This wouldn't make insurance a perfect regulator. Insurers are accountable to their own balance sheets, not to the public as a whole, and they may respond to poorly understood risks by excluding them rather than making them safer, as has been the case with some types of weather disaster in areas that are becoming more prone to things like flooding and wildfires.But insurance companies do have to convert uncertainty into prices, contractual language, and technical requirements, which makes some currently difficult to quantify things more quantifiable, at least monetarily.The AI incidents this summer caused no reported material damage, which is one reason they're getting relatively little coverage, despite being fairly meaningful events. The insurance industry sees these sorts of narratives through the lens of cost and risk, though, and this is a category of loss with no conventional attacker, no stolen credential, several plausible defendants, and almost no claims history, arriving at a moment in which companies are racing to give autonomous systems more access to all of their systems—a lot of valuable and potentially vulnerable infrastructure.The people whose job is to price that risk haven't decided what it costs, yet. And until they do, what they add to or remove from their application forms may be more consequential to the norms and expectations in this space than what the government mandates, on the matter.Show Noteshttps://www.businessinsurance.com/as-ai-agents-go-rogue-cyber-insurers-are-adapting-their-policies/https://www.investing.com/news/stock-market-news/as-ai-agents-go-rogue-cyber-insurers-are-adapting-their-policies-4878768https://openai.com/index/hugging-face-model-evaluation-security-incident/https://openai.com/index/hugging-face-incident-and-the-road-ahead/https://huggingface.co/blog/security-incident-july-2026https://huggingface.co/blog/agent-intrusion-technical-timelinehttps://www.anthropic.com/news/investigating-incidents-cybersecurity-evalshttps://cyberunit.com/insights/ai-sandbox-escapes-three-labs-meta-anthropic-openai/https://labs.cloudsecurityalliance.org/research/csa-research-note-frontier-ai-models-hacking-real-systems-ev/https://techcrunch.com/2026/07/07/the-worst-hacks-and-breaches-of-2026-so-far/https://blackkite.com/reports/2026-ransomware-reporthttps://www.cisa.gov/news-events/cybersecurity-advisories/aa23-320ahttps://www.munichre.com/en/insights/cyber/cyber-insurance-risks-and-trends-2026.htmlhttps://www.swissre.com/risk-knowledge/advancing-societal-benefits-digitalisation/about-cyber-insurance-market.htmlhttps://www.marsh.com/en-gb/services/international-placement-services/insights/global-insurance-market-index.htmlhttps://compyl.com/guides/cyber-insurance-readiness-guide/https://www.aon.com/en/insights/articles/cyber-and-tech-e-and-o-market-reporthttps://www.cybersecuritydive.com/news/merck-settlement-notpetya-insurance/703922/https://therecord.media/mondelez-and-zurich-reach-settlement-in-notpetya-cyberattack-insurance-suithttps://assets.lloyds.com/media/eb6de9ce-293b-4213-80f8-9dc69c45b1a9/Y5381%20Market%20Bulletin%20-%20Cyber-attack%20exclusions.pdfhttps://www.whitehouse.gov/wp-content/uploads/2026/06/eo-14409.pdfhttps://www.axios.com/2026/08/04/inside-trump-ai-framework This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit letsknowthings.substack.com/subscribe
Topics covered in this episode: Pandas Should Go Extinct Pydantic-pint puts real-world units in your Pydantic models How Libraries Run Rust Inside Python (With PyO3) AWS acquires DuckLabs Extras Joke Watch on YouTube Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Pandas Should Go Extinct Pandas' slowness pushes teams toward "Big Data" tools (Spark, Databricks) they don't actually need — most workloads never hit true Big Data scale Amazon Redshift telemetry: ~95% of tables are under 100GB, ~87% of queries touch 80GB or less — that's "Medium Data," not Big Data Polars and DuckDB fill that gap: single-machine, fast, no cluster required 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory On a real-world NYC taxi dataset (3GB parquet), pure DuckDB ran 2x faster than pure Pandas while using a fraction of the RAM Bonus: Apache Arrow lets you pass data between Pandas/Polars/DuckDB with zero copying, so trying them out doesn't mean a full rewrite Michael #2: Pydantic-pint puts real-world units in your Pydantic models Pydantic-pint bridges Pydantic and Pint so models can validate physical quantities like 4m or 12 meters instead of bare floats. Fields annotated with PydanticPintQuantity parse user input, convert between compatible units, and serialize quantities back out as strings. That closes a real gap for anything consuming API payloads, config files, or sensor data with measurements, letting you enforce units at the validation boundary instead of hoping every caller remembered them. via PyCoder's Weekly newsletter Unit mix-ups have literally crashed spacecraft; now your Pydantic models can refuse them at the door. Annotate a field as Annotated[Quantity, PydanticPintQuantity('km')] and inputs like 12 meters arrive auto-converted to kilometers Validation covers string, numeric, and quantity inputs, and model_dump_json serializes quantities as readable unit strings Installable from PyPI as pydantic-pint, MIT licensed, with docs at pydantic-pint.readthedocs.io Early-stage solo project at version 0.4, so API stability and maintenance are open questions worth discussing Calvin #3: How Libraries Run Rust Inside Python (With PyO3) Pydantic v2's validation core (pydantic-core) is Rust under the hood, built with PyO3 — this post shows how that bridge actually works via a small hand-built JSON parser Four steps to get Rust into Python: write a normal Rust module, annotate with PyO3 macros (#[pyfunction], #[pymodule]), compile/install with maturin, then just import it The parser builds a Rust tree first — Python never touches it until the boundary crossing Key insight: converting the Rust result into Python objects (.into_pyobject) is often the expensive part, not the parsing — 100,000 JSON values means ~100,000 Python objects built after parsing's already done Errors cross the boundary too: Rust's typed errors convert into real Python exceptions (ValueError, FileNotFoundError) via From/?, so callers get clean Python semantics Takeaway for anyone porting Rust in: if you're returning a scalar, don't sweat it; if you're returning a big structure, profile the boundary — that's the real cost, not the algorithm Michael #4: AWS acquires DuckLabs Thank you Dylan McConnell. What does this mean for the DuckDB ecosystem? DuckDB is the open-source in-process analytical SQL engine. MIT licensed. The IP is not owned by any company - it's held by the nonprofit DuckDB Foundation, which was created when the team spun out of CWI Amsterdam. Peter Boncz, the CWI representative on the Foundation board, describes it as the entity that holds all IP of open-source DuckDB. DuckLabs (ducklabs.com) is the company, formerly branded DuckDB Labs. Founded a little over five years ago by Hannes Mühleisen and Mark Raasveldt to give the DuckDB team a stable long-term home, bootstrapped deliberately instead of taking VC, grown to 30+ people in Amsterdam, funded by support and feature-prioritization contracts. It employs the core devs. It does not own DuckDB. DuckLake is one of three projects DuckLabs builds, what they call the Duck Stack: DuckDB, DuckLake, and Quack. DuckLake is the lakehouse format that puts catalog metadata in a SQL database instead of in files on object storage. Quack is newer - an RPC-style protocol that turns DuckDB into a client-server system where both ends are DuckDB instances, slated to stabilize in DuckDB v2.0 in September 2026. MotherDuck is a separate Seattle company, Jordan Tigani's, selling serverless hosted DuckDB. It was started in partnership with DuckDB Labs and has worked closely with Hannes and Mark for four years. It contracted DuckLabs for engineering work and contributes heavily upstream - three of its engineers are among the top 10 outside contributors to DuckDB. It also sells its own DuckLake offering. Customer and collaborator, never owner. What the AWS post changes. Amazon bought the company, not the project. DuckLabs joined AWS effective September 1, with the process concluding August 31, 2026. Hannes and Mark keep leading the team and the project's technical direction, the team stays in Amsterdam, and DuckDB stays MIT under the Foundation. AWS gets the people and a direct line to the roadmap. The license protects your code, not your priorities. Three second-order effects worth tracking: The Foundation board is the real question. It has three directors: Mühleisen, Raasveldt, and Boncz. Two now work for AWS. Commentary on the deal has focused on exactly this - the license protects the code, not the roadmap. The announced counterweight is governance: a technical advisory board on the Foundation, and opening the extension stack so extensions signed by other developers can run in DuckDB. MotherDuck immediately moved into the business DuckLabs vacated. It now sells DuckDB enterprise support, which it had avoided because it didn't want to compete with DuckLabs' business model, and says it has explicit blessing from Hannes and Mark now that they're joining Amazon. It also bought Tower.dev the day before the AWS announcement. Everyone expects an AWS DuckDB service. Tigani says Amazon will likely release one eventually, and welcomes the competition, citing Redshift's failure to slow Snowflake on AWS. The groundwork is already visible: Amazon Quick uses DuckDB to query S3 Tables and has processed over 2.5B queries with it since launching in October 2025. The DuckLake angle is the one to watch. AWS is heavily committed to Iceberg through S3 Tables, and it just acquired the team behind a competing lakehouse format. The stated plan is to use DuckDB, DuckLake, and Quack together to power a new generation of data services, but which format wins internal priority is unannounced. Extras Calvin: astral-sh/uv 0.12.12: code-signed release binaries
Send us Fan MailEvery click, purchase, transaction and customer interaction creates data—but who turns those numbers into useful business decisions?In this episode of The Kapeel Gupta Career PodShow, explore the career of a Data Analyst, one of the most versatile data-driven career options today.Discover what data analysts actually do, the career scope in India and abroad, essential skills, educational pathways, tools such as Excel, SQL, Python and Power BI, salary potential, top colleges and career opportunities.If you enjoy numbers, patterns, problem-solving and technology, this episode will help you understand whether Data Analytics could be the right career for you.
Are your social feeds filling up with automated slop and made-up stories? In this episode of Digital Marketing From The Coalface, Dave and Julie return from their summer break to tackle the fine line between smart AI productivity and cringey, low-quality automation. If you're an engineering company leader looking for genuine online visibility without the fluff, this episode shows you how to put AI to work to provide high-value data analytics while keeping human strategy firmly in the driver's seat. What we cover in this episode: The cringey AI lead magnet test: Julie shares what happened when she tested a tool that promised 30 "tailored" LinkedIn posts, only to receive completely fake case studies and bizarre claims. Talking to your data in plain English: Dave and Julie talk about how connecting tools like Windsor.ai to Google Search Console, Google Ads, and GA4 lets you query your marketing data in natural language, bypassing complex PowerBI filters and SQL queries. AI watermarks & editorial control: AI watermarking isn't something to fear if you take ownership of your content. You need to make sure to apply human oversight, and refine every word you're sharing. Strategy vs. "random acts of marketing": tech businesses need a structured marketing framework so that their teams aren't just creating ad-hoc AI content for the sake of it. Plus, Julie and Dave swap summer highlights and office drama, featuring Edinburgh Fringe comedy, metal detecting adventures, a new coffee machine, and some spilled milk that no one cried over. Enjoyed the episode? Browse more at redevolution.com/digital-marketing-from-the-coalface
"Realnie nie robisz review kodu. Przy tej ilości nie masz opcji, żeby to zrobić."
Talk Python To Me - Python conversations for passionate developers
How many files does your query read before it reads any data? On some data lakes, you go through JSON and metadata files first, just to learn which Parquet files matter. DuckLake asks one SQL question instead. The metadata lives in a real database. The data stays in plain Parquet. That's the entire format. Pedro Holanda joined DuckDB in 2018, when it was still a research prototype at CWI. He's the lead DuckLake developer. Guillermo Sanchez Dionis works on DuckLake and the new Quack protocol. With Quack as the catalog, DuckLake handles 200 transactions a second under heavy contention. No other open table format comes close. Episode sponsors Six Feet Up Talk Python Courses Links from the show Guests Pedro Holanda: pedroholanda.org Guillermo Sanchez: linkedin.com PhD on progressive indexes: ir.cwi.nl SQLite: www.sqlite.org Litestream: litestream.io boring hardware: talkpython.fm DuckDB: duckdb.org episode 491: talkpython.fm Iceberg: iceberg.apache.org manifesto: ducklake.select DuckLake: ducklake.select spec: ducklake.select this diagram: blobs.talkpython.fm Data inlining: ducklake.select ducklake-dataframe: github.com Polars course: training.talkpython.fm CSV parser: duckdb.org Zero-copy Arrow: duckdb.org ART index: duckdb.org async I/O: duckdb.org v1.0: ducklake.select Git-like branching: ducklake.select Watch this episode on YouTube: youtube.com Episode #562 deep-dive: talkpython.fm/562 Episode transcripts: talkpython.fm Theme Song: Developer Rap
Die letzten Monate waren geprägt von neuen Foundation Models für tabellarische und sequentielle Daten: TabPFN 3 skaliert auf eine Million Beobachtungen, Google stellt mit TabFM ein eigenes Modell samt BigQuery-Integration vor, NXAI veröffentlicht TiRex-2 mit Kovariaten-Unterstützung, und an der Spitze von GIFT-Eval steht mit STRIDE eine Kombination aus LLM-Reasoning und Time Series Foundation Model. Dazu kommen der ClickHouse-MCP-Server und der Stand der Umsetzung des EU AI Act nach dem Digital Omnibus. Im Praxisteil vergleichen wir TabICL v2 mit einem getunten XGBoost, Meta Prophet und naiven Baselines auf stündlichen NO2-Messwerten von fünf Messstationen, ausgewertet über ein Jahr rollierender Kreuzvalidierung mit dem Mean Absolute Scaled Error. Wir zeigen, welches Feature-Engineering nötig ist, wie sich der Vorteil von TabICL mit der Länge der Historie verändert und was das an Rechenzeit kostet. Zum Schluss ordnen wir ein, wann sich ein Foundation Model für Zeitreihen anbietet und wann XGBoost die pragmatischere Wahl bleibt. **Zusammenfassung** TabPFN 3 (Mai 2026) skaliert auf einer H100 auf bis zu 1 Mio. Beobachtungen; verbessertes KV-Caching senkt die Prognosezeit auf 0,1–3 ms pro Testbeobachtung und macht das Modell für schnelle Batch-Prognosen nutzbar. Googles TabFM ist von TabPFN und TabICL inspiriert, liegt im TabArena-Benchmark vor TabPFN 3 und ist direkt in BigQuery integriert. TiRex-2 von NXAI setzt auf eine xLSTM- statt Transformer-Architektur und kann jetzt zusätzliche Kovariate einbeziehen – ein Test steht bei uns noch aus. STRIDE führt den GIFT-Eval-Benchmark an: Das LLM prognostiziert nicht selbst, sondern steuert über destillierte Embeddings ein Time Series Foundation Model. Kurz notiert: Der ClickHouse-MCP-Server (v0.4.1) erlaubt LLM-Abfragen ohne SQL, etwa zur Log-Diagnose; beim EU AI Act gelten die Transparenzpflichten seit August, die Kennzeichnung von Bestandssystemen greift ab dem 2.12.2026. Praxis-Setup: TabICL v2 mit Kalender-, Fourier- und Lag-Features gegen getuntes XGBoost, Prophet, TabICL out-of-the-box sowie Naive und Seasonal-Naive; stündliche NO2-Daten, 24-Stunden-Horizont, Metrik MASE. Ergebnisse: Bei zwei Jahren Historie liegt TabICL klar vorn, bei rund 8.000–9.000 Trainingsbeobachtungen ist XGBoost praktisch gleichauf, bei drei Monaten Historie noch etwa 4 % schlechter; ohne jedes Feature-Engineering schlägt TabICL Prophet und die naiven Baselines deutlich. Kosten: Die Kreuzvalidierung mit TabICL auf einer L40S-GPU dauert etwa 17-mal länger als mit XGBoost, auf CPU ist das Modell nicht praktikabel – Caching dürfte diesen Nachteil künftig verkleinern. **Links** Link zum begeleitenden Blogartikel "TabICL v2 für Zeitreihen: Das In-Context-Learning-Modell im Vergleich mit XGBoost und Meta's Prophet" https://www.inwt-statistics.de/blog/tabicl_v2_fuer_zeitreihen #72: TabPFN: Die KI-Revolution für tabulare Daten mit Noah Hollmann https://www.podbean.com/ew/pb-94ri2-18aca83 #57: Mehr als heiße Luft: unsere Berliner Luftschadstoffprognose mit Dr. Andreas Kerschbaumer https://www.podbean.com/ew/pb-u6xwt-16ff139 TabPFN-3 Technical Report: https://priorlabs.ai/technical-reports/tabpfn-3 TabPFN auf GitHub: https://github.com/PriorLabs/TabPFN Google Research zu TabFM: https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/ TabFM in BigQuery: https://cloud.google.com/blog/products/data-analytics/tabfm-adds-predictive-ml-to-bigquery TiRex-2 (NXAI): https://www.nx-ai.com/en/tirex-2 | Code: https://github.com/NX-AI/tirex-2 | Paper: https://arxiv.org/abs/2607.01204 STRIDE – Reasoning-Aware Training for Time Series Forecasting: https://arxiv.org/abs/2605.08625 Time Series LLMs am Beispiel t0-alpha: https://towardsdatascience.com/time-series-llms-explained-with-t0-alpha/ ClickHouse MCP Server: https://github.com/ClickHouse/mcp-clickhouse EU AI Act nach dem Digital Omnibus (Überblick): https://www.deloitte.com/de/de/issues/innovation-ai/eu-ai-act-digital-omnibus.html TabICL v2 auf GitHub: https://github.com/soda-inria/tabicl TabICL-Dokumentation: https://tabicl.readthedocs.io/en/latest/ Tutorial zum TabICLForecaster: https://tabicl.readthedocs.io/en/latest/tutorials/time_series_forecasting.html Meta Prophet: https://github.com/facebook/prophet GIFT-Eval Leaderboard: https://huggingface.co/spaces/Salesforce/GIFT-Eval TabArena Leaderboard: https://huggingface.co/spaces/TabArena/leaderboard
Today's Black Tech Building Show. Two major discussions topics. First, What is GCC. Lastly, looking at the latest trends in SQL and it relations to AI. Finally, the latest tech news.Recorded 7/28/2026
Guy and Eitan discuss an interesting case study where crazy high CPU was detected after an upgrade to SQL Server 2025. Relevant links: SQL 2025 showing crazy-high CPU | LobsterPot Solutions Enable Automatic Tuning - Azure SQL Database & Azure SQL Managed Instance | Microsoft Learn Plan Forcing in SQL Server - Erin Stellato
In this episode, Ray Cochrane breaks down Hot Chips 2026, the engineering conference where IBM, NVIDIA, Intel, AMD, Arm, and Fujitsu all showed how their next processors actually work. The headline disclosure is a mainframe core that runs Arm natively. Ray also covers Apple’s odd M6 Mac mini naming, London’s first autonomous Uber rides, Amazon’s purchase of the company behind DuckDB, GitHub’s HydraFusion, the best of IFA 2026, and new USDA research on farmed salmon. – Want to start a podcast? It’s easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens the show with a personal update. He apologizes for the late rollout and the missed Monday episode, having spent the week fighting a cold. He’s also heading to Michigan to spend time with family and visit his father’s gravesite. He hopes to record a couple of shows from his dad’s old studio while he is there, including a special episode planned for Tuesday. He also points listeners to a fresh site redesign that trades the old techie look for something cleaner and friendlier. The featured segment starts from a wrap-up post on Arm’s newsroom. However, Cochrane broadens it to cover the whole conference rather than a single article. Hot Chips has run every August since 1989, and this year’s event was the 38th, held August 23rd through the 25th at Stanford’s Memorial Auditorium. What Hot Chips Actually Is Cochrane draws a line between Hot Chips and the big consumer trade shows. CES and Computex exist for product announcements and marketing. Meanwhile, Hot Chips is an IEEE engineering conference where chip architects present block diagrams and die photos for thirty minutes at a stretch. The audience matters as much as the content. Roughly five hundred people who design chips for a living fill the room, and they would spot a fudged number immediately. No written paper is required, just the talk and the slides. In-person tickets sold out this year, as did Stanford’s dorm housing. Cochrane says he plans to cover the conference annually going forward. IBM Built a Mainframe Core That Speaks Arm The disclosure that stopped Cochrane cold came from IBM. Its next processor for IBM Z and LinuxONE runs two completely different instruction sets natively, on every one of its eleven cores. Those are z/Architecture, IBM’s own mainframe language, and AArch64, which is 64-bit Arm. Crucially, this is not emulation. Nor is it Arm cores glued onto the die beside the mainframe cores. IBM built 2,792 Arm instructions directly into the hardware, which it says is more than double the mainframe instruction count. Each core carries two separate decoders while sharing the caches, branch predictor and register files downstream. It switches between the two in nanoseconds. IBM even added dedicated hardware to flip byte order, because Arm and the mainframe store numbers in opposite directions. The payoff is Arm SystemReady compliance, meaning off-the-shelf Arm Linux runs on a mainframe unmodified. Patrick Kennedy of ServeTheHome, who was in the room, wrote: “I am sitting here still in awe of what IBM is doing here; this is not Z plus Arm cores, this is Z and Arm in one core.” Cochrane flags one precision point that is easy to get backward. IBM did not license Arm’s core designs and drop them in. Instead, it took its own mainframe core and taught it AArch64 under an architecture license, which is considerably harder engineering. The specifications are striking. The chip uses a 2nm process, with eleven cores running above 5.7GHz sustained and no turbo mode at all. Each core gets 36MB of L2 cache, backed by a 3.5GB virtual L4 pool. Furthermore, the reliability target is eight nines, which works out to roughly three tenths of one second of unplanned downtime per year. IBM gave it no name and no ship date, though the press expects “Telum III” around 2028. Arm, Fujitsu and NVIDIA Show Their Hands Arm itself had plenty to discuss, starting with an unfortunate name. Its first chip in thirty-five years is called the AGI CPU, which is a product name rather than any claim about artificial general intelligence. For three and a half decades, Arm designed processor blueprints and licensed them out, collecting royalties without competing. That era is now over. The AGI CPU is Arm’s own silicon, co-designed with lead customer Meta, running up to 136 cores on TSMC’s 3nm process at 300 watts. Arm’s CEO says the company has more than $2 billion in customer demand across the next two fiscal years. Fujitsu brought the detail Cochrane called the coolest of the conference. Its MONAKA chip packs 144 Arm-based cores, but the trick is the cache. Rather than sitting alongside the cores and eating die area, the entire last-level cache lives on a separate 5nm die with the 2nm compute die stacked directly on top. It ships in 2027 in 350-watt and 500-watt versions. NVIDIA had more stage time than anyone, with six sessions. Its new Vera CPU carries 88 cores of NVIDIA’s own Olympus design, which marks a change: the previous Grace CPU used Arm’s off-the-shelf cores. Consequently, Arm’s win here is the instruction set, not the blueprint. The memory disclosure drew the most attention, with a fully loaded system reaching 1.5TB at 1.2TB/s while the whole memory subsystem draws just 30 to 40 watts. The Caveat on NVIDIA’s Benchmark Slides Cochrane pushes back on how NVIDIA presented its numbers. On the standard SPEC integer benchmark, Vera scored 925 against AMD’s 128-core EPYC score of 898, about three percent ahead. However, the slide NVIDIA showed normalizes that same result per physical core, which makes a three percent gap look enormous. NVIDIA defends the choice, arguing that per-core throughput matters when thousands of AI agents run at once. Cochrane grants that it is a fair argument to make. Even so, his verdict is blunt: it is a different number from the headline one, and presenting it that way is not the best look. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. Apple’s Newest Chip Landed in Its Cheapest Mac Last week’s episode covered the Mac Studio half of Apple’s August announcement. Tonight Cochrane takes the other half, the new Mac mini, and finds the numbering genuinely strange. The $899 base Mac mini gets the M6, which is Apple’s first 2nm chip and the newest process the company has shipped. It brings twelve CPU cores, twelve GPU cores, and 170GB/s of memory bandwidth. Apple also introduced a third CPU core class called super cores. So Apple’s most advanced chip sits in its cheapest desktop, while the M5 Pro, M5 Max and M5 Ultra above it all carry a lower number. It gets stranger. The M6 mini has Thunderbolt 4 while the pricier M5 Pro mini has Thunderbolt 5, and the memory ceiling runs backward too. Apple explains none of it across three announcement pages. Reading between the lines, Cochrane figures the M6 is the entry point of a new generation that shipped ahead of its larger siblings. London’s Robotaxis Started Carrying Passengers Arm’s monthly roundup covers everything outside the data center, and Wayve stood out. Transport for London granted the British self-driving company private hire vehicle licenses on August 5th, the same category a minicab needs. Then on September 3rd the service launched with Uber, marking the first autonomous rides ever offered to UK passengers. A small fleet of Ford Mustang Mach-Es covers anywhere in London except the airports, with a TfL-licensed safety driver still aboard. Over 140,000 Londoners signed up. The compute runs on NVIDIA’s Arm-based automotive platform, which is why Arm claims the win. Cochrane notes the pattern: once you set a standard nobody can move off, these wins keep arriving. Two more items round it out, both about squeezing AI onto phones. Google’s Pixel 11 shipped in August with the Tensor G6, and Google claims on-device AI runs up to 3.5x faster using 3.5x less energy. Those are Google’s own unbenchmarked figures. Separately, Graphcore’s research team worked with Arm to run an 11-billion-parameter vision model on phone-class processors by squeezing each parameter to 2.7 bits, taking the model from roughly 22GB down to 3.7GB. Intel Is Pitching AI Infrastructure From a Long Way Back Intel’s newsroom post previews the AI Infra Summit, running September 15th through 17th in Santa Clara. CEO Lip-Bu Tan takes a fireside chat on Tuesday morning, with three Intel sessions across the show. Cochrane unpacks two terms first. Physical AI means AI that acts in the real world through sensors and motors rather than living on a screen, and Intel takes it seriously enough to have renamed its PC division the Client Computing and Physical AI Group in May. Disaggregated inference splits the two phases of running a model: reading your prompt is compute-hungry, while writing the answer back is memory-hungry. The context is where this gets interesting. Intel’s revenue rose 25 percent last quarter, its fastest growth since 2011. Nevertheless, in AI accelerators the company barely registers next to NVIDIA. Gaudi has effectively been abandoned, to the point that Intel stopped maintaining its open-source driver, and AMD passed Intel in data center revenue last quarter. Tellingly, when Intel demoed disaggregated inference at Computex, NVIDIA GPUs handled one phase and SambaNova chips the other while Intel supplied the coordinating CPU. Crescent Island, Intel’s actual inference chip, does not sample until later this year, and Intel declined to publish its memory bandwidth. Amazon Bought the Company Behind DuckDB On August 26th, Amazon signed a deal to acquire DuckLabs, the Amsterdam company behind DuckDB. Cochrane spends time explaining what DuckDB is, since listeners outside the data world may never have encountered it. The problem it solves is familiar. Querying a large pile of data files traditionally meant either running a database server or spinning up a data warehouse with a cluster, a bill, and a loading pipeline. Both are heavy machinery for a question you wanted answered in ten seconds. DuckDB instead ships as a library rather than a server. You add it to your program like any other package, point it at your files, and write ordinary SQL directly against them. The data never moves. The usual comparison is SQLite, which is embedded in nearly every phone and browser on earth. Where SQLite excels at looking up one record, DuckDB rebuilds that embedded idea for chewing through millions of rows. It now sees roughly 62 million monthly downloads on Python’s package index alone, up from about 25 million last October. AWS says it is buying the company, not the project. DuckDB stays free and open source under the MIT license, held by a Dutch nonprofit foundation, and the founders join AWS while continuing to run technical direction from Amsterdam. Andy Warfield, a VP and distinguished engineer at AWS, described DuckDB as “the glibc of structured data: a lean, unglamorous, ubiquitous dependency that a great deal of software links against and almost nobody has to think about.” Cochrane sits with what that ownership means. He reaches for an analogy: imagine Daniel Stenberg selling curl. He doubts it would ever happen, but the concept alone is startling given how much infrastructure depends on it. His read is that AWS is betting DuckDB becomes as foundational as curl and SQLite already are. GitHub Has One AI Model Grade Another One’s Homework GitHub shipped Project HydraFusion into Copilot as a research preview. Instead of routing your request to a single model, it picks one of three approaches per request. Sometimes one model simply answers. Alternatively, a cheaper model drafts, and a quality gate decides whether to escalate. The interesting one is Critique. One model writes the code, a separate model from a different family reviews it read-only, and the original gets one pass to revise. Despite the name, nothing is fused here. There is no voting and no merging, just one model at a time with a gate deciding whether to spend more. GitHub explained the reasoning in an earlier post: “a model reviewing its own work is still bounded by its own training biases: the same training data and techniques, the same blind spots.” Research supports it. A team at NeurIPS in 2024 showed that models recognize their own writing and score it higher than human graders do. GitHub’s earlier number had a Claude Sonnet and GPT critic pairing closing about three-quarters of the gap between Sonnet and the larger Opus model. This mirrors a workflow Cochrane uses constantly and has described on a previous episode. He runs a cross-check review with a second model from a different company, and it routinely surfaces issues the first model missed. He explains that different training data, different engineers, and different reinforcement approaches build different internal biases about what counts as correct. Looking ahead, he expects more of these “Frankenstein patterns” where models from different training families work together. The Best of IFA 2026, and What You Can Actually Buy IFA opened to the public in Berlin for its 102nd year, with about 1,900 brands. Cochrane splits The Verge’s roundup in two, since much of what generates headlines at these shows never ships. Starting with real products, iRobot’s flagship Roomba Max 875 Combo runs $1,199 and ships in about two weeks. Its SealForce feature drops a hidden skirt from the chassis when it detects carpet, sealing against the fibers so suction concentrates instead of leaking out the sides. That reaches 35,000 pascals, iRobot’s strongest yet. A step-down model at $899 carries the same trick, though Cochrane balks at both prices. Anker’s Soundcore Sleep 4 Pro earbuds arrive in November at $349.99. The charging case carries its own round touchscreen, so you pick soundscapes, set alarms, and read sleep stats without your phone. Optical sensors read heart rate and variability from the ear canal, which beats the wrist for accuracy, and the case masks a snoring partner. Cochrane remains unconvinced about sleeping with earbuds in. Philips also has smart rope lights, the Hue Liane 360, on sale now. They glow evenly around the tube rather than showing individual LEDs. They also run $400 for three meters, which works out to about $130 per meter of rope light. As for concepts nobody can buy, iRobot showed a robot vacuum that carries a smaller robot vacuum on its back in a garage and lowers it to deploy. Lenovo brought a 14-inch laptop whose screen rolls out to 17 inches at the press of a button, which reviewers call the first rollable that feels close to shippable. Tecno showed a phone with essentially no border around the screen, and Acer had a Windows gaming handheld that swivels its screen up over a keyboard. Two themes ran through the show. Humanoid robots were the loudest thing on the floor, and IFA’s own CEO framed the event as being about robots that work rather than robots that demo. Meanwhile, AI stopped being its own product category and became an ingredient, showing up in refrigerators, treadmills, dishwashers, and motorized TV mounts. Farmed Salmon Isn’t the Omega-3 Machine It Used to Be USDA scientists measured farmed salmon and found considerably less of the good fat than the government’s own database claims. EPA and DHA are two fatty acids you get almost entirely from fish. Your body can build them from the plant version, but only in tiny amounts, so the NIH’s position is that eating them is the only practical way to raise your levels. Those fatty acids are structural pieces of every cell, with DHA concentrating in the brain and retina. That is why the federal dietary guidelines, issued jointly by USDA and Health and Human Services, recommend at least eight ounces of fish a week and steer you toward salmon. That amount is calibrated to deliver about 250 milligrams a day. Published in Frontiers in Nutrition last month, the study found EPA and DHA in farmed Atlantic salmon came in 54.7 percent lower than USDA’s own reference values, last updated in 2018. A three-ounce serving fell from roughly 1,670 milligrams to about 756. Consequently, two servings a week now fall about 14 percent short of the target. Plant-derived fats meanwhile rose two to three times over. The likely cause is feed. Salmon are carnivores, and farms once fed them oily little fish. There was never going to be enough of those as the industry scaled, so crops filled the gap: soy, canola, sunflower and linseed. Importantly, the study does not claim to have proven this and calls the feed shift a plausible explanation. Independent corroboration lends it credibility. Researchers at Stirling measured a similar halving in Scottish salmon between 2006 and 2015, and Norway’s marine institute saw it across thousands of samples. There is a land dimension too. Roughly half the world’s soy grows in South America, where rainforest gets cleared for feed. Matthew Hayek, who studies the environmental cost of protein at NYU, told Inside Climate News that “soy is a major, important protein and oil ingredient in fish farming.” Adding up two decades of soy across all fish farming, he puts the extra forest clearing at around the area of Nicaragua or Bangladesh. That figure covers all fish farming rather than salmon alone, and Hayek notes it is hard to attribute soy use to any single species. Cochrane closes with two caveats. First, the study measured only eight fish, bought around Maryland, DC and Virginia over six weeks in 2023, and nearly all sourced from Chile. That is not a national survey, and the authors say plainly the sample was not large enough to change government advice. Rather, it flags that a federal database value needs rechecking, which is what the paper set out to do. Second, on whether you should care, farmed salmon still beats beef, chicken and eggs by a mile, since those carry essentially zero EPA and DHA. What it loses is its crown among fatty fish, dropping to mid-pack behind herring, sardines and mackerel and roughly level with trout. The broader health case is also softer than the 2000s suggested. A review of 86 trials covering 162,000 people found supplements barely moved heart attacks or deaths, so eating fish and swallowing fish oil are not the same claim. No producer has responded to the findings, and USDA, whose own scientists ran the study, declined an interview and did not answer emailed questions. Cochrane wraps up with housekeeping and a note that he will be back on Labor Day. The post The Mainframe Learned to Speak Arm #1875 appeared first on Geek News Central.
Liquid Weekly Podcast: Shopify Developers Talking Shopify Development
In this episode of the Liquid Weekly Podcast, hosts Karl Meisterheim and Taylor Page are joined by Alexander Hupfer, founder of Fusion Metrics and a solo Shopify app developer.The conversation digs deep into the current state of Shopify app analytics following Mantle's shutdown, why certain metrics are far more subjective than they appear, and how AI agents are changing the way developers query their own data.Sponsor - The Support HeroesSTAY CONNECTEDSubscribe to Liquid Weekly for more expert insights: https://liquidweekly.com/EPISODE HIGHLIGHTS- Alexander's Origin Story- Why Shopify Doesn't Build This- The Mantle Shutdown- Metrics That Aren't Black and White- Installs as a Vanity Metric- AI Agents + Analytics: Using MCP servers and agents like Claude Code and Codex to write SQL queries - Build vs. Buy vs. Self-HostABOUT ALEXANDER HUPFERAlexander is the founder of Fusionmetrics, an analytics platform built specifically for Shopify apps. Coming from a theoretical physics background, he moved into the analytics space in 2022 and pioneered install attribution by linking Google Analytics events with Partner API data.FIND ALEXANDER ONLINE & RESOURCES- Fusion Metrics: https://fusionmetrics.com/- Alexander's Twitter/X: https://x.com/alexanderhupfer- PartnerMetrics by Björn Forsberg (open source): https://github.com/forsbergplustwo/partner-metrics- Matt De Sousa / Blair (ex-Mantle) conversation on attribution: https://orange-aphid-848.notion.site/Live-with-Blair-Beckwith-Shopify-Mantle-Growth-the-Future-of-Shopify-Apps-3c15075fb6d680a5bd14d8ec230430e4- Alexander's talk from last year's Mantle event on churn modeling: https://youtu.be/9g7N7VdHfWU?si=48yb1k9xM6sb9VwX- Darius's Dev Dashboard revamp screenshots: https://x.com/darius_gai/status/2087916085794283800?s=20- Alexander on Ilias' Podcast talking analytics: https://youtu.be/eC8i9ex6xnQ?si=97o59ZK31c-SfCgfTIMESTAMPS00:00 - Cold open: The mask-selling business00:30 - Introduction & the bat/rabies saga02:20 - Alexander's origin story: COVID masks to Shopify apps06:42 - Moving into analytics & the physics background11:03 - Current state of Shopify app metrics & the Mantle shutdown13:49 - Why Shopify doesn't provide revenue analytics (the ROI question)16:43 - The Dev Dashboard revamp & transaction vs. MRR21:57 - Reconstructing MRR from the Partner API event stream25:11 - Installs as a vanity metric & search ranking27:28 - Churn and LTV: the most complicated metrics34:11 - App acquisition & assessing performance objectively36:40 - The three Rs: Acquisition, Retention, Revenue37:06 - Attribution & app store conversion rates38:28 - Using AI agents & MCP to query analytics43:01 - GA4 attribution on the App Store vs. storefront45:31 - App Store optimization & keyword-stuffed titles48:17 - Build vs. Buy vs. Self-Host53:17 - Picks of the Week56:35 - Dev ChangelogDEV CHANGELOG- Oxygen is now available on development stores: https://shopify.dev/changelog/oxygen-is-now-available-on-development-stores- WebMCP support for Liquid and Hydrogen storefronts: https://shopify.dev/changelog/webmcp-liquid-hydrogen- Standard storefront events and actions now support cart attributes: https://shopify.dev/changelog/events-and-actions-cart-attributes-support- Hydrogen developer preview update, August 18, 2026: https://shopify.dev/changelog/hydrogen-developer-preview-update-august-18-2026- App intents on admin.app.intent.link now open as a full-page navigation: https://shopify.dev/changelog/app-intents-on-admin-app-intent-link-now-open-as-a-full-page-navigationPICKS OF THE WEEK- Karl: "Star City" on Apple TV — a sister show to "For All Mankind" exploring an alternate history where Russia reached the moon first- Alexander: "The Capitalist Manifesto" by Johan Norberg- Taylor: "Project Hail Mary" (the movie) — looking forward to reading the book next
Two teams pitch the same figure for ‘active customers' right before a leadership meeting. One dashboard displays 40,000 while the other dashboard displays 52,000. No mistakes were detected by the teams; however, the difference in the stats is due to one team or employee deploying a different data filter and the other using another join path.Now there's an AI agent involved, who is simply asked to 'retrieve the most recent number of active customers'. Not realising that there are two versions, it selects one, states that figure with complete confidence, and then goes about its business. In that moment, the semantic gap ceases to be merely an internal discussion and becomes a business risk.On this episode of the Don't Panic! It's Just Data podcast, host Shubhangi Dua, Podcast Producer and B2B Tech Journalist at EM360Tech, is joined by Ananya Devraju, a Business Intelligence Developer at Premier Inn, to discuss the semantic gap and how AI is rendering an old data problem. They look at the situation in which metric definitions are spread across five different dashboards rather than kept in a single central location, and explain why an AI agent can produce a perfectly valid SQL query yet return a completely incorrect answer.The problem, consistent even with AI in the picture, is governance. AI has just made the issue of governance more apparent. “It was never a data quality problem, but it's a governance of meaning problem, and it's been sitting there long before AI showed up. AI has just made it louder now,” Devraju tells Dua.Devraju will be presenting a talk at Big Data LDN (BDL) this year on Thursday, September 24, on When Migrations Break Your Metrics: Rebuilding Data for Commercial Analytics. The talk is scheduled for 3:20 pm to 3:50 pm in the Data Architecture Modernisation Theatre.TakeawaysSemantic gaps existed before AI; AI has only made them louder and faster.The same metric is displayed on different dashboards, each using different date logic, filters, or join paths.There can be five different definitions of an "active user" across the five dashboards, all of which are incompatible.SQL can be syntactically correct and yet provide the wrong answer to a business question.Validation should be placed at the definition level, not at the SQL level.AI agents can query the raw staging tables rather than referring to the governed marts by name alone.Incorrect joints or the wrong grain result in numbers that look plausible but are actually wrong.According to dbt Labs research, 72 per cent of peopleprioritise AI-assisted coding while 24 per cent prioritise validation.Incorrect data is more dangerous than having no data; it is the confident errors that pose a risk.Since end users are unable to check the answers at the time they occur, trust has to be established earlier on.Provide a semantic layer that is governed and has version control, with designated owners for each metric.Ad hoc SQL statements on no side-channel should generate the "official" figures that are reported.Chapters00:00 Introduction to the Semantic Gap in AI02:56 Understanding Data Interpretation and Governance05:56 The Importance of Centralized Definitions09:05 AI Hallucinations and Data Validity12:02 Navigating Governance in AI14:56 The Role of AI in Data Validation17:56 Final Thoughts on AI and Data MeaningGuestAnanya Devraju — Business Intelligence Developer, Premier InnHostShubhangi Dua — Podcast Producer & B2B Tech Journalist, EM360TechListen to more episodes of Don't Panic! It's Just Data on EM360Tech.#SemanticGap #AIAgents #DataGovernance #DataAnalytics #SemanticLayer #AIData #BusinessIntelligence #DataEngineering #EnterpriseAI #DataQuality
Nathan's guest this episode is Pete Johnson, Field CTO of AI at MongoDB, and the conversation is really two conversations woven together: a history of database architecture, and a status report on the still-unsolved problem of agent memory. Pete opens with a framing device that recurs throughout — he was born in February 1970, four months before E.F. Codd's original relational-model paper that gave rise to SQL. The relational model, he explains, was built for a world where storage was the scarce resource, so normalization — splitting data across linked tables to avoid duplication — was the rational design choice. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/write-change-recall-forget-mongodb-s-pete-johnson-on-how-retrieval-drives-agent-performance/ Sponsors: Mercury: Mercury is the banking platform loved by 300,000+ entrepreneurs, with virtual cards and Spend controls for granular budgets, receipts, and low-risk AI agent purchases. Learn more and apply in minutes at https://mercury.com Granola: Granola is an AI-powered notepad that securely transcribes meetings and turns rough notes into clean, structured action items. Try it free at https://granola.ai/tcr Diffusion: Diffusion helps organizations build custom AI software factories that scale business outcomes, not just outputs. Cognitive Revolution listeners get a 25% service credit on their first engagement at https://diffusion.io/tcr Deepgram Flux TTS: Deepgram Flux TTS brings lifelike AI voices with real personalities that handle interruptions, pauses, and natural conversation. Try all the voices free through September 12 at https://deepgram.com/keep-talking Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr CHAPTERS: (00:00) About the Episode (03:15) Sponsor: Mercury (04:56) SQL versus NoSQL (11:24) Enterprise database choices (Part 1) (18:30) Sponsors: Granola | Diffusion (21:27) Enterprise database choices (Part 2) (21:27) Schema flexible search (33:50) Contextualized chunking tradeoffs (Part 1) (35:12) Sponsors: Deepgram Flux TTS | Claude (37:17) Contextualized chunking tradeoffs (Part 2) (46:22) Retrieval quality thresholds (53:17) Agent memory systems (01:05:51) Enterprise AI deployment (01:16:32) Voyage acquisition strategy (01:23:38) Global AI adoption (01:30:56) Episode Outro (01:34:41) Outro PRODUCED BY: https://aipodcast.ing
Si llevas años usando Bash, Zsh o Fish y piensas que los pipes de Unix son lo más parecido a la perfección, este episodio te va a hacer tambalear los cimientos. Porque existe un shell que no pasa texto entre comandos: pasa estructuras de datos. Tablas, listas, registros, fechas, tamaños de archivo con tipo real. Y encima habla con Ollama sin que tengas que escribir ni una línea de Python.Ese shell es Nushell. Está escrito en Rust, tiene más de 40.000 estrellas en GitHub, y su filosofía es sencilla: los pipes deberían transportar datos con tipo, no texto que luego parseas con awk, sed o jq.En este episodio te cuento mi experiencia pasando de Fish a Nushell con ejemplos reales. Cuando escribes ls no obtienes texto: obtienes una tabla con columnas tipadas. Puedes hacer ls | where size > 1mb | sort-by size sin recurrir a awk ni números mágicos. El shell entiende qué es un filesize, qué es una fecha, qué es un número.Y luego está open, que entiende el formato por la extensión: JSON, YAML, TOML, CSV, SQLite... todo se convierte en datos estructurados. Abres un SQLite y ejecutas consultas con query db. Y todo combinable: http get a una API, filtrar con where y guardar con save — en un solo pipeline, sin archivos temporales.La guinda es la integración con IA. Como Nushell entiende JSON y Ollama habla JSON, se entienden a la perfección. Te enseño un pipeline que lista procesos, filtra los que consumen más de 100MB de RAM, se los manda a un modelo local, y mata el que más memoria usa. Todo en una línea. También te hablo de ai.nu, un módulo que envuelve Ollama, OpenAI y DeepSeek, con function calling desde el shell.También hago una comparativa: Bash, Zsh, Fish y Nushell cara a cara. Bash funciona en cualquier sitio pero el manejo de datos es arcaico. Zsh es Bash con esteroides pero los pipes siguen siendo texto. Fish es moderno pero no entiende de tipos. Nu es el único con estructuras de datos de verdad. PowerShell fue el primero en pasar objetos, pero Nu es lo que PowerShell debería haber sido.Capítulos del episodio:0:00 — Introducción: de Bash a Fish, la evolución de las shells2:30 — El problema del texto plano: por qué Nushell es diferente5:00 — La trifecta: ls, where y select, SQL en tu terminal7:30 — Tipos reales: la shell entiende fechas, tamaños y números10:00 — Open: abrir JSON, CSV, YAML y SQLite sin herramientas externas13:00 — Procesamiento avanzado: $in, save, append y par-each15:30 — HTTP GET: APIs de GitHub y meteorología desde la shell18:00 — Comparativa de shells: Bash vs ZSH vs Fish vs Nushell21:00 — Nushell e IA: integración nativa con Ollama sin Python24:00 — Instalación, casos de uso y conclusiones finalesMás información y enlaces en las notas del episodio
In this week's Data Debrief, the companion show to Driven by Data: The Podcast, Kyle Winterbottom and Catherine Dowden-King unpack the week's main episode with Michael Ross and range far beyond it into the collapse in graduate hiring, the succession planning nobody is doing, and what's really happening at both ends of the data job market.From a record 45% drop in advertised graduate roles, to the experienced leaders who've been out of work for two years, to Michael's case that every average hides an opportunity, Kyle and Catherine make the argument that AI is taking the blame for decisions plenty of businesses already wanted to make, and that the bill for not developing people will land in about five years' time.They also discuss:Why a 45% drop in advertised graduate jobs is the lowest figure ever recorded, and why AI can't be held responsible for all of it.How record university enrolment colliding with a shrinking entry-level market creates a problem unfolding in real time.Why "entry-level" data roles asking for two years of Python or SQL were never really entry-level.What happens to the pipeline when the admin-heavy tasks juniors cut their teeth on get absorbed by agents.Why the real risk isn't AI replacing juniors, but having nobody ready when the current workforce retires.How data roles are shifting towards QA, product management and facing back into the business.Why succession planning has only ever been pointed at the top of the house, and why that has to change.What skills matrices and career pathways expose the moment you ask "and when this bottom layer moves up, then what?"Why some organisations announced AI-driven headcount cuts when the business was simply performing badly.How "we're cutting because of AI" got turned into a PR positive rather than a negative.Why a retailer, a telco and an airline sat at the same table are nowhere near the same stage of the journey.What the senior end of the market actually looks like, and why it gets discussed far less than the graduate end.Why there are more head of, director and VP roles than at any point in fifteen years, even as true CDO roles decline.How being overqualified has become as much of a barrier as being underqualified.Why an entire cohort of data leaders has been tarred with the same brush through no fault of their own.How the failure to prove value from data and analytics now has a direct, downstream human cost.What Michael Ross's epiphany moment says about technical specialists becoming commercial operators.Why de-averaging matters more than any dashboard, and how averages quietly mislead entire teams.How an 80% average occupancy hid the fact that no hotel was anywhere near 80%.Why 100% occupancy might be a pricing failure rather than a success story.What it takes for a CEO to get close enough to the commercial detail of their own business to win.Why putting your head above the parapet takes bravery, and why the cost of not doing it is the situation the industry is now in.Why Dolly Parton's Imagination Library may be the most important thing she ever built.What's left of the Future of Data, AI & BI event, Driven by Data Live on 8 October at Tobacco Dock, and the new roles on the NED Appointment Finder.
Topics covered in this episode: Web UIs for your reverse proxy Wagtail 8.0 is hot off the presses RISC-V is now officially supported by CPython Django's annual releases make every version an LTS Extras Joke Watch on YouTube About the show Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Michael #1: Web UIs for your reverse proxy Traefik, nginx, and Caddy all sit in front of a lot of self-hosted infrastructure, and all three are configured by hand-editing files. Three active projects put a control plane on top: Traefik Manager (Python + Flask), Nginx UI (Go + Vue), and caddy/ui (React + Node). All three are additive rather than replacements - none of them take ownership of your config away from you - which is the part that matters when the thing has write access to production routing. Traefik Manager is the Python one: Flask 3.1 and Gunicorn for the control plane, a lightweight Go agent for remote instances, currently v1.10.0 with an Android companion app. Nginx UI is a single Go binary at 11.3k stars, with a block-style config editor, an Ace editor doing LLM completion on nginx syntax, and an MCP server so agents can drive it. caddy/ui runs as two containers next to your existing Caddy, reads and writes your Caddyfile directly, and uses Caddy's /adapt API to validate before reload - no Docker socket required. Each one edits the config the underlying server already reads, so your files stay the source of truth and you can drop the UI without unwinding anything. Undo is a first-class feature across all three - timestamped backups with optional Git history, config version compare and restore, Caddyfile snapshots with one-click rollback. Observability is where they diverge: Traefik Manager does CrowdSec and a visual route map, Nginx UI does server metrics, caddy/ui streams access logs over SSE and pulls p50/p95/p99 off Caddy's Prometheus endpoint. Maturity spread is wide - Nginx UI has 11.3k stars, caddy/ui has 4 and was built in a single Claude session - and caddy/ui ships with auth off by default, so set CADDY_UI_USER and JWT_SECRET before it goes anywhere near a public interface. Calvin #2: Wagtail 8.0 is hot off the presses Link: https://github.com/wagtail/wagtail/releases/tag/v8.0 Custom base page models are now supported, so projects aren't locked into subclassing Wagtail's Page as shipped (Matt Westcott). New v3 REST API handles both read and write CMS operations, a first for Wagtail's API. A global registry for permission policies, plus full customizability for the remaining page views via PageViewSet. AVIF and WebP images are no longer auto-converted to PNG by default, a real behavior change to watch on upgrade. Five security fixes: page admin API restrictions, document identification by SHA1 hash, descendant collections in the Documents/Images API, snippet copy permissions, and the page translation endpoint. Formalized Django 6.1 support, and CI now runs on uv with a lockfile. Sponsor: Logfire from Pydantic Your AI agent failed at 2am. Was it the model? A tool call? The database? Most observability tools can't tell you, because they only see part of your stack. Pydantic Logfire sees all of it. One trace across your agents, LLMs, APIs, and database. Down to the infrastructure: services, Kubernetes, and hosts. It's built on OpenTelemetry, with SDKs for Python, TypeScript, and Rust, and it works with any OTel-compatible language. Every prompt, token count, and cost, right next to your vector searches and API calls. You query everything with Postgres-compatible SQL. And so can your coding agent, through the Logfire MCP server. Stop guessing. Read the trace. Pydantic Logfire. AI, it's still just engineering. Visit pythonbytes.fm/logfire today and sign up today. Get 10M records free every month, no card required. You can even click “Onboard with your coding agent” to copy a prompt to have claude or codex integrate Logfire into your app. Thanks to Pydantic for supporting the show. Calvin #3: RISC-V is now officially supported by CPython Link: https://blog.python.org/2026/08/riscv-now-officially-supported/ CPython added RISC-V as a tier 3 platform under PEP 11, specifically the 64-bit Linux target riscv64-unknown-linux-gnu. RISC-V is an open ISA anyone can implement, unlike x86 and ARM, and its market is projected to quadruple by 2032. The RISE Project donated real RISC-V machines for buildbots; the author's work was funded by a Sovereign Tech Agency fellowship. What changes: the port is now a maintained compatibility target, so CPython changes are less likely to quietly break it. What doesn't: no python.org installers, no binary wheel parity for native extensions. Next up: RISC-V runners in CPython CI for pre-merge feedback, then a push toward tier 2, plus architecture-specific optimizations. The ask is testing. If you have RISC-V hardware, build CPython, run your test suite, file what breaks. Tier 3 is the weakest support tier. PEP 11 tier 3 requires a core developer contact and a buildbot, but failures on tier 3 platforms explicitly do not block a release. Saying "ongoing CI/testing expectations" oversells it. The honest bit is "someone is now on the hook for it, and breakage gets noticed," not "it's guaranteed working." Worth the caveat that this is Linux SBCs, not microcontrollers. A VisionFive 2 counts, an ESP32-C6 or Pico 2 does not. Those are 32-bit non-Linux parts where MicroPython is still the answer. Michael #4: Django's annual releases make every version an LTS Starting with Django 2028, Django will move to one January feature release per year, adopt calendar-based version numbers, and support every release for three years. The old distinction between standard and LTS releases disappears, giving teams a predictable annual upgrade path that aligns more closely with Python's own release and support cadence. Every Django release becomes the safe, long-supported choice, so teams no longer need to wait for a specially designated LTS version or absorb two years of changes at once. Each release gets one year of mainstream bug fixes followed by two years of security and data-loss fixes. New releases support the three latest Python versions and add the next Python release during their first year. Calendar versioning begins with Django 2028, followed by Django 2029 and so on. Three Django versions will be supported at any time, giving third-party packages a clearer rolling target. Nothing changes before 2028, and existing commitments for Django 5.2 LTS and 6.2 LTS remain in place. Extras Calvin: The Python docs now document the time complexity of built-in types https://docs.python.org/3.16/library/time-complexity.html Thinking in Python - Bruce Eckel's free book https://thinkinginpython.com/ Michael: prune_uv_pythons.py - Prune uv-managed Python installs, keeping only the newest patch per minor version Runs automatically in my system “upgrade” script: upgrade-output-2026.png Started using Ollama cloud models for my Hermes assistant. Thanks to Jeff Triplett I learned they are not just local models. Joke: The Tao of Programming - Book Seven: Corporate Wisdom
This episode is a re-air of one of our most popular conversations, featuring insights worth revisiting. This week on The Data Stack Show, Eric and welcomes back Ruben Burdin, Founder and CEO of Stacksync as they together dismantle the myths surrounding zero-copy ETL and traditional data integration methods. Ruben reveals the complex challenges of two-way syncing between enterprise systems like Salesforce, HubSpot, and NetSuite, highlighting how existing tools often create more problems than solutions. He also introduces Stacksync's innovative approach, which uses real-time SQL-based synchronization to simplify data integration, reduce maintenance overhead, and enable more efficient operational workflows. The conversation exposes the limitations of current data transfer techniques and offers a glimpse into a more declarative, flexible approach to managing enterprise data across multiple systems. You won't want to miss it. Highlights from this week's conversation include: The Pain of Two-Way Sync and Early Integration Challenges (2:01) Zero Copy ETL: Hype vs. Reality (3:50) Data Definitions and System Complexity (7:39) Limitations of Out-of-the-Box Integrations (9:35) The CSV File: The Original Two-Way Sync (11:18) Stacksync's Approach and Capabilities (12:21) Zero Copy ETL: Technical and Business Barriers (14:22) Data Sharing, Clean Rooms, and Marketing Myths (18:40) The Reliable Loop: ETL, Transform, Reverse ETL (27:08) Business Logic Fragmentation and Maintenance (33:43) Simplifying Architecture with Real-Time Two-Way Sync (35:14) Operational Use Case: HubSpot, Salesforce, and Snowflake (39:10) Filtering, Triggers, and Real-Time Workflows (45:38) Complex Use Case: Salesforce to NetSuite with Data Discrepancies (48:56) Declarative Logic and Debugging with SQL (54:54) Connecting with Ruben and Parting Thoughts (57:58) The Data Stack Show is a weekly podcast powered by RudderStack, customer data infrastructure that enables you to deliver real-time customer event data everywhere it's needed to power smarter decisions and better customer experiences. Each week, we'll talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data. RudderStack helps businesses make the most out of their customer data while ensuring data privacy and security. To learn more about RudderStack visit rudderstack.com. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Building something people actually want is supposed to be the happy ending. But it arrives with a bill attached: feature requests you didn't ask for, pull requests you'd rather not maintain forever, users demanding the one thing you swore you'd never build, and — if you're unlucky — a company where the sales team quietly starts deciding what engineering works on. DuckDB has spent the last two and a half years working through that list. So how do you stay a database engineering team when success keeps trying to turn you into something else?Hannes Mühleisen, co-creator of DuckDB, is back to talk through the answers they've landed on. Their fix for unwanted pull requests was an extension mechanism, which then forced them to make every part of the engine pluggable — including the parser, which meant ripping out 20,000 lines of Postgres' yacc grammar and rewriting SQL parsing on top of PEG. Their fix for the users demanding client-server was Quack, a protocol designed by people who'd already published a paper on why every existing database wire protocol is wrong. And their answer to Apache Iceberg, after three years of implementing it, was DuckLake: throw out the Avro-and-JSON metadata files and keep the metadata in a database, on the grounds that the Iceberg REST catalog has a Postgres in it anyway.Which brings us to the news Hannes breaks in this episode: DuckDB Labs is being acquired by AWS, while the DuckDB Foundation, the project and its licence stay where they are. There's the question of why a profitable, self-funded, 30-person company in Amsterdam would take that deal, what commitments you write into the contracts when you're worried today's promises might outlive today's management, and what it's actually like to have a boss again after five years without one. If you're curious how an open source project keeps its technical soul once the enterprise arrives — or you just want to know why parsing SQL is harder than parsing almost anything else — Hannes has some good answers.---Support Developer Voices on Patreon: https://patreon.com/DeveloperVoicesSupport Developer Voices on YouTube: https://www.youtube.com/@DeveloperVoices/joinOur previous episode with Hannes: https://youtu.be/pZV9FvdKmLcDuckDB: https://duckdb.org/DuckDB Foundation: https://duckdb.foundation/DuckLabs (formerly DuckDB Labs): https://ducklabs.com/DuckLake: https://ducklake.select/Quack (DuckDB's client-server protocol): https://duckdb.org/quack/DuckDB v2.0: Your Database Deserves a Better Parser: https://duckdb.org/2026/08/20/duckdb-20-peg-parserRuntime-Extensible Parsers (CIDR 2025 paper): https://duckdb.org/pdf/CIDR2025-muehleisen-raasveldt-extensible-parsers.pdfDon't Hold My Data Hostage (VLDB 2017 paper): https://www.vldb.org/pvldb/vol10/p1022-muehleisen.pdfcpp-peglib: https://github.com/yhirose/cpp-peglibGNU Bison: https://www.gnu.org/software/bison/PEP 617 – New PEG parser for CPython: https://peps.python.org/pep-0617/PRQL: https://prql-lang.org/Apache Iceberg: https://iceberg.apache.org/CWI (Centrum Wiskunde & Informatica): https://www.cwi.nl/en/DuckCon #7, Amsterdam: https://duckdb.org/events/2026/06/24/duckcon7/Kris on Bluesky: https://bsky.app/profile/krisajenkins.bsky.socialKris on Mastodon: http://mastodon.social/@krisajenkinsKris on LinkedIn: https://www.linkedin.com/in/krisjenkins/
Help us become the #1 Data Podcast by leaving a rating & review! We are 67 reviews away! Real recruiter spends 20 seconds on this resume and finds nothing worth keeping. I show you why.
Manu Mehra is Head of AMER Industries Strategic Deal Pricing at Databricks, with more than 12 years of experience across pricing, product, cloud, and AI, including Google Cloud and Thermo Fisher Scientific. He brings a practical perspective on how AI is changing the way companies think about outcomes, value, platforms, and pricing models. In this episode, Manu explains why traditional pricing models don't neatly fit AI, why outcome-based pricing is compelling but difficult to standardize, and how companies can turn platforms into solutions around specific business problems. Mark challenges him throughout the conversation, especially on the attribution problem: if AI creates the value, how do you know AI actually caused it? Why You Have to Check Out Today's Podcast: Learn why AI is pushing pricing toward outcomes. Discover how platforms become solutions customers will pay more for. Understand the attribution and standardization challenges behind AI pricing. "Pricing cannot be an afterthought. It has to be integrated within the product roadmap." — Manu Mehra Topics Covered: 01:15 – How an accidental pricing analytics role led Manu to a career spanning product, cloud, AI, sales, finance, and strategic deal pricing. 03:30 – Why Pricing Has to Start With the Product. Why integrating pricing into the product roadmap can create value before launch instead of scrambling for cost-plus pricing afterward. 05:30 – Why AI Breaks Traditional Pricing Models. Why subscription, license, and consumption models don't fully fit AI when thousands of customers can pursue completely different outcomes 08:00 – The Hardest Problem With Outcome-Based Pricing. Why AI outcomes are difficult to standardize across billing, finance, legal, and revenue recognition—and why 10,000 customers could mean 10,000 different outcomes 11:00 – What Actually Counts as an AI Outcome? Manu uses a QBR example where AI can automate 95% of the SQL work, turning hours and effort saved into a measurable form of value. 13:30 – The Attribution Problem: Did AI Really Create the Value? Mark and Manu debate how to determine whether AI actually caused increased revenue, lower costs, or other gains—or simply helped the business get there faster. 16:00 – Platform vs. Solution: What Are You Really Selling? Why a broad platform can have wildly different value depending on the customer's use case—and how platforms can become solutions by solving specific business problems. 19:00 – How to Turn Products Into Business Solutions. Manu explains how compute, data, and AI layers can be combined into packaged solutions instead of being sold as isolated products. 21:30 – How Customer Segmentation Makes AI Pricing Scalable. Why identifying recurring customer patterns can help companies map different business problems to repeatable combinations of SKUs instead of creating a custom solution for every customer. 24:00 – Why AI Companies Use Credits. How credits can create cost predictability, manage backend costs, and give customers flexibility across different AI capabilities. 27:00 – When Credits Make Sense—and When They Don't. Why platform customers may value the flexibility of credits while digital-native customers who already know exactly what they want may have less need for them. 30:00 – The Pricing Advice Manu Wants Leaders to Hear. Why pricing should never be an afterthought and why the industry is moving from cost-plus toward value-based and outcome-based pricing. Key Takeaways: "The reason is, even though you might be a platform organization or you're selling a platform, but end of the day, you're still trying to solve a customer problem." — Manu Mehra "The tricky thing with outcome is it's very hard to standardize it." — Manu Mehra "Pricing needs to be integrated during the product roadmap." — Manu Mehra Connect with Manu Mehra: LinkedIn: https://www.linkedin.com/in/manumehra1/ Connect with Mark Stiving: LinkedIn: https://www.linkedin.com/in/stiving/ Email: mark@impactpricing.com
What happens when a 40-year-old legal data company decides its employees should start building their own software? This week on we talk with Best Lawyers CEO Phillip Greer and Senior Vice President of Research and Product Strategy Elizabeth Petit about an internal AI transformation that reaches far beyond adding ChatGPT to the corporate toolkit. Best Lawyers is experimenting with generative engine optimization, internal agentic systems, vibe coding, and an AI development environment where employees across research, finance, marketing, and other departments build applications around the company's data.Greer begins with a challenge facing every law firm marketing team: traditional search is changing. Google AI Overviews and answer engines such as ChatGPT, Claude, and Gemini increasingly give users answers without sending them to the familiar list of blue links. Greer argues that SEO still matters, but law firms now need to think about Generative Engine Optimization, or GEO, and the signals AI systems use when deciding which sources deserve trust. Structured data, schema markup, substantive content, and third-party validation all become part of the equation. For Best Lawyers, its long history of peer-reviewed rankings offers an interesting advantage. The company's data serves as an independent signal that AI systems might weigh differently from content produced by a firm's own marketing department.Petit explains how Best Lawyers is applying the same thinking to legal marketing through Smithy AI, a system designed to help attorneys and law firm marketers develop profile content without endlessly copying the same biography across websites. Smithy draws from Best Lawyers' structured information and existing lawyer content to produce a starting point that attorneys and marketers then edit. The larger goal is authenticity. As generative systems make producing generic legal content almost effortless, Greer argues that distinctive expertise, voice, and credible third-party signals become more valuable rather than less.The conversation then moves inside Best Lawyers, where Greer has taken a far more unusual approach to AI adoption. After building a secure data layer connecting systems including SQL databases, HubSpot, Gong, Google Analytics, and accounting data, he created an internal Best Lawyers App Store where employees use natural language to build applications against company data. What began with roughly 30 percent of the workforce vibe coding has grown to around 40 percent, according to Greer. Petit describes building research and KPI dashboards despite coming from a research rather than software engineering background. Projects that once required Excel formulas, Power BI reports, development queues, and weeks of waiting now sometimes move from a question at 9:30 to a working internal application by 10:30.That shift also changes the role of professional software engineers. Rather than spending their time building another reporting screen or internal form, Best Lawyers' engineers increasingly concentrate on architecture, data infrastructure, performance, governance, and the guardrails surrounding employee-built applications. Greer describes moving parts of the company's data architecture toward Elasticsearch and developing “Bestie,” an internal agentic AI team member. Yet speed introduces another problem. Petit and Greer describe an “AI vampire” effect, where instant feedback encourages people to keep working because the machine never gets tired, goes home, or stops responding. Human judgment includes knowing when the human needs to stop.Listen on mobile platforms: Apple Podcasts | Spotify | YouTube | Substack[Special Thanks to Legal Technology Hub for their sponsoring this episode.]Email: geekinreviewpodcast@gmail.comMusic: Jerry David DeCicca
https://clearmeasure.com/developers/forums/ Anna Hoffman is a Principal Group Product Manager for Azure Data on Microsoft's SQL Engineering team, where she works across SQL Server, Azure SQL, and SQL database in Fabric. She is also the host of Data Exposed, Microsoft's video series covering the data platform. Anna holds an engineering degree from Georgia Tech along with a master's degree in data science, and started her career as a data and applied scientist before moving into product management. She continues to blog regularly for the Azure SQL Devs' Corner. You can find her on GitHub and on X at @AnalyticAnna. Anna's Devblogs - https://devblogs.microsoft.com/azure-sql/author/antho/ Anna's Github - https://github.com/amthomas46 X Account - https://x.com/analyticanna LinkedIn - https://www.linkedin.com/in/amthomas46/ What's New Across Microsoft Microsoft SQL Server Blog Author Page Azure SQL Devs' Corner Author Page Data Exposed (YouTube playlist, 500 episodes) Azure SQL YouTube channel "The Era of the Agentic Database Developer" (Build 2026 recap) "What's new across Microsoft SQL in 2026 so far" (mid-year recap) MSSQL extension roadmap - https://github.com/microsoft/vscode-mssql/wiki/roadmap GitHub Copilot Agent Mode in SSMS - https://learn.microsoft.com/ssms/github-copilot/agent-mode Previous Appearances on the Azure & DevOps Podcast: Episode 160 - https://azuredevopspodcast.clear-measure.com/azure-sql-database-with-anna-hoffman-episode-160 Want to Learn More? Visit AzureDevOps.Show for show notes and additional episodes.
Nik and Michael are joined by Shaun Thomas to discuss estimating work_mem, memory management in general, and writing an extension to help. Here are some links to things they mentioned: Shaun Thomas https://postgres.fm/people/shaun-thomaswork_mem https://www.postgresql.org/docs/current/runtime-config-resource.html#GUC-WORK-MEMpg_stat_database https://www.postgresql.org/docs/current/monitoring-stats.html#MONITORING-PG-STAT-DATABASE-VIEWALTER ROLE SET configuration_parameter https://www.postgresql.org/docs/current/sql-alterrole.html#SQL-ALTERROLE-PARAMS-CONFIGURATION-PARAMETERhash_mem_multiplier https://www.postgresql.org/docs/current/runtime-config-resource.html#GUC-HASH-MEM-MULTIPLIERRecent Postgres releases with 28 CVEs fixed https://www.postgresql.org/about/news/postgresql-186-1711-1615-1519-1424-and-19-beta-3-released-3365/Systemic Risks in the Managed PostgreSQL Industry (part 1 of 6, Mehmet Ince) https://mehmetince.net/part-1-6-systemic-risks-in-the-managed-postgresql-industry-extension-risks-are-real-exploiting-postgis-memory-corruption-bug-at-neondb-supabase-and-many-more/Improving Postgres Connection Scalability: Snapshots (blog post by Andres Freund) https://techcommunity.microsoft.com/blog/adforpostgresql/improving-postgres-connection-scalability-snapshots/1806462~~~What did you like or not like? What should we discuss next time? Let us know via a YouTube comment, on social media, or by commenting on our Google doc!~~~Postgres FM is produced by:Michael Christofides, founder of pgMustardNikolay Samokhvalov, founder of Postgres.aiWith credit to:Jessie Draws for the elephant artwork
„Nie ma czegoś takiego jak dostatecznie dobre rozwiązanie. Zawsze można lepiej, zawsze można głębiej.” Co się dzieje, gdy z pozycji prezesa w jednej z największych instytucji finansowych świata przechodzisz do organizacji, która działa w tempie startupu i od pierwszego dnia każe kwestionować status quo?
PHP Podcast – August 20, 2026 Hosts: Joe Ferguson, Sara Golemon & Holly Schilling Shirley MacLaine is the answer — but what’s the question? RAM prices went up 500%, GitHub fell over for eight hours, and the crew figured out how you can actually contribute to PHP. Glasses optional. Shirley MacLaine, Green Room Shade, and the Show’s New Normal The episode opens with Joe flustered — not because it’s a “flustered day,” but because there’s shade being thrown in the green room backstage chat. Rather than spoil the drama, the crew turns a mysterious green-room answer into a running bit: “Shirley MacLaine” is the answer, and listeners are invited to submit what the question was. It’s declared the first round of PHP Architect’s Jeopardy, complete with its own Easter egg sound cue. Joe also lays out where the show is heading. Eric and John took over the podcast the previous week to relive their glory days, and the plan is for the old guys to step back to roughly one episode a month. Joe and Sara are working on something to fill one slot, another mystery show is in the works pending signed contracts, and Alive and Kicking is confirmed still alive and kicking after a great recent episode with Derek. Joe also plugs the PHP 8.6 Beta 1 tag and Scott Keck-Warren’s PHP Community Podcast interview with release managers Daniel Scherzer and Matteo Bacotti. The RAM Apocalypse: 500% in Twelve Months The big topic of the week is the ongoing memory and chip crisis. Memory prices have climbed 500% in twelve months, and the crew commiserates about the parts they wish they’d bought bigger. Sara explains the brutal math: you can’t build a semiconductor fab fast enough — two or three years minimum — and by the time one comes online nobody knows if we’ll have overproduction or a burst bubble. Worse, Sara notes that essentially every RAM stick through the end of 2027 is already accounted for and sold to vendors. The ripple effects are everywhere. Console prices are going up instead of down late in their cycle, with next-gen consoles projected to start at $1,000 or more. The hobbyist single-board computer market is getting crushed — Pine64 announced they’re stepping back from making hardware, and a Raspberry Pi 5 that debuted cheap now runs over 100 pounds. Holly’s new PC, bought at the start of 2025, ended up shipped 3,000 kilometers the wrong way to California and is stuck awaiting import paperwork, leaving her leaning on a laptop and some now-precious Raspberry Pi zeros. The nostalgia gets thick as the crew reminisces about DIP chips, EDO RAM, Pentiums, Epson 486s, Tandy 1000s, and the TRS-80. Special venom is reserved for Packard Bell (“utter freaking trash”) and Gateway 2000’s cow-pattern branding — which Sara confirms sold better in Wisconsin than anywhere, though hay bales in a Berkeley storefront still boggle the mind. A chat comment about technicians de-soldering and reballing BGA RAM chips becoming economically viable gets a hearty “absolutely.” Memory-Aware Development and Why PHP 7 Doubled Down Bringing the RAM crisis back to PHP, Joe wishes more developers were aware of how their applications consume and release memory — a lesson he credits to learning enough C back in the day, where you have to manage memory yourself. He connects sloppy memory awareness to the N+1 query problems web developers keep tripping over. Sara drops a great deep dive: a significant reason PHP 7 was roughly twice as fast as PHP 5 was changes in the memory layout. Every variable became referenced by one fewer pointer, and while eight bytes sounds trivial, every level of indirection adds time across every single instruction and access. Sara adds the CPU-level detail — one layer of indirection can be a single instruction on most architectures, but adding a second layer can push a lookup from one instruction to three. That leads into a warm tangent about learning C to become game developers. Sara’s evergreen joke: “I’m going to be a game developer” is the programmer’s version of “I’m going to buy a bar.” Great people, brutal hours, endless competition, and the reality of hitting spacebar 400 times to figure out why you can phase through a wall. The cat-reading-the-paper “I should buy a boat” meme makes an appearance to seal it. The GitHub Outage and the Monoculture Problem Monday’s eight-hour GitHub outage hit the crew directly. Holly couldn’t use a site that only offered “log in with GitHub,” and Joe got kicked out of his CLI auth session mid-PR with no way to re-authenticate. To GitHub’s credit, they published an incident update and a follow-up blog post: a service auto-scaled so aggressively to handle network traffic that the sidecar and supporting services couldn’t keep up, bringing the whole thing down. The conversation turns to whether this is self-inflicted. Joe recalls GitHub’s pre-Microsoft, gold-standard reliability and wonders aloud how much the decline lines up with Copilot’s arrival and internal AI adoption, with uptime reportedly slipping below a single nine at points. Sara defends them somewhat — the number of actions, CPU cores, and pull requests has genuinely hockey-sticked, partly because AI has emboldened people who previously wouldn’t have opened a PR. But as Joe puts it, the call is coming from inside the house, since GitHub itself has been pushing AI. On alternatives, Joe says the least-jarring migration for PHP Architect’s clients would be self-hosting GitLab, since GitHub Actions and GitLab runners are nearly identical in syntax — Atlassian’s Bitbucket, by contrast, is a bridge too far, mostly because the entire ecosystem assumes you’re on GitHub. Sara names this the core problem: monoculture. The crew discusses package mirrors, local caches, 12-factor thinking, and Composer’s support for custom mirrors, all while remembering the PHP repo intrusion years ago that came from an unmaintained self-hosted Git server. The takeaway: owning your pipeline end-to-end is the only way an outage can’t stop you — and Joe teases spinning up a self-hosted GitLab now that “the boss” (Sara) has signed off. How to Contribute to PHP (and Handling Security Reports) The crew highlights two PHP Foundation blog posts. First, Matt Stauffer’s “How to Contribute to PHP,” adapted from a talk he gave at Atlanta PHP. It goes well beyond “learn C,” clearly separating the PHP project, the PHP ecosystem, and the Foundation, and lays out approachable on-ramps: testing pre-releases (PHP 8.6 Beta 1 is out, Beta 2 lands next week), improving documentation, and writing tests — which, spoiler, are written in PHP, not C, using PHP’s own test format that’s simple enough to learn from any single example. Other contribution paths include triaging and reviewing issues across PHP repositories — invaluable work that frees core developers from wading through AI-generated slop bug reports — and participating in internals via the well-documented mailing list process, up to and including running for release manager (8.7 managers will be needed before you know it). Sara points folks to discord.phpc.chat for the PHP Discord, with dedicated Internals and Foundation channels for anyone the mailing list intimidates. Second, Sebastian Bergmann’s “So you received a security report. Now what?” is a jump-around reference for application developers rather than a front-to-back read, walking through roughly ten steps to triage, validate, and resolve reported issues the right way. Sara shares a real-world example from mobile: a flagged package that was only exploitable on a rooted device with an actively hostile package installed alongside it — a very different risk profile than a SQL injection on an API endpoint. Cue reminiscing about writing SQL against Access databases over ODBC from PHP (and Perl) back in the 90s, and Joe’s advice for handling any security report: don’t panic, and always know where your towel is. Links from the show: PHP Tek 2027 — April 27–29, 2027 in Chicago; early bird tickets & hotel available now PHP Tek 2027 CFP Audio versions of the podcast at phparch.com Join us live on Discord at discord.phparch.com PHP Discord — discord.phpc.chat Community Corner Podcast: PHP 8.5 + 8.6 Release Manager Daniel Scherzer Memory prices climb 500% in 12 months So You Received a Security Report. Now What? How to Contribute to PHP Shirley MacLaine Host: Joe Ferguson Mastodon: @joepferguson@phpc.social PHPArch.me: @svpernova09 Sara Golemon Mastodon: @pollita@phpc.social Holly Schilling Mastodon: @TheCodeLorax@tech.lgbt Streams: Youtube Channel Twitch Connect & Hire PHP Architect Website Twitter/X Mastodon Hire PHP Developers Looking to hire PHP developers? Email support@phparch.com – Joe and the team are available for consulting, infrastructure work, Ansible playbooks, and code review. Partner This podcast is made a little better thanks to our partners Displace Infrastructure Management, Simplified Automate Kubernetes deployments across any cloud provider or bare metal with a single command. Deploy, manage, and scale your infrastructure with ease. https://displace.tech/ OurCVEs Your security posture, on autopilot with OurCVEs CodeRabbit Cut code review time & bugs in half instantly with CodeRabbit. PHP Architect Consulting Your PHP codebase deserves a partner, not a contractor PHP Architect provides long-term technical partnerships for organizations that need senior-level PHP expertise that you can depend on https://www.phparch.com/consulting/ Music Provided by Epidemic Sound https://www.epidemicsound.com/ Join Us Live Next Week Youtube Channel Got feedback? Join us on Discord at discord.phparch.com The post The PHP Podcast 2026.08.20 appeared first on PHP Architect.
Will AI agents solve the data industry's hardest challenges, or are we just amplifying existing chaos?In this episode, I chat with Oliver Laslett, Co-founder and CTO at Lightdash, to explore the frontier of agentic analytics, data modeling, and modern software engineering practices. We dive deep into how AI is shifting daily workflows from writing SQL and DBT models to higher-order system thinking, why the hardest data problems remain human and organizational, and the cultural differences in tech optimism between London and San Francisco.Oliver also breaks down why permissions, curation, and data lineage are still the true bottlenecks, and why high-ownership communication matters more than ever in an era of AI slop.
Topics covered in this episode: Python 3.12.14, 3.11.16, 3.10.21 - security releases Codeberg's AI-code ban tests its role as a GitHub alternative Brett Cannon: what's missing for reproducible builds on PyPI nothing records the source code a distribution came from. direct_url.json captures it when you install from a repo or archive, so the fix is putting the same info in sdist/wheel metadata. recording the build tools. Wheels can already do this via PEP 770 SBOMs in .dist-info/sboms/ - sdists can't, since they're a tarball plus a precalculated PKG-INFO with nowhere to hang extra metadata. Either "don't use sdists" or an sdist v2. Extra extra extra, hear all about it Extras Joke Watch on YouTube Sponsored by Logfire from Pydantic pythonbytes.fm/logfire This episode is brought to you by Pydantic Logfire. It's observability for AI apps from the team behind Pydantic - agents, LLMs, APIs, database, and infrastructure in a single trace, queried with Postgres-compatible SQL. Your coding agent can query it too, through their MCP server. I'll tell you more later. Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Python 3.12.14, 3.11.16, 3.10.21 - security releases https://blog.python.org/2026/08/python-31214-31116-31021/ Source-only security releases for the three branches now in security-fix-only mode; release team blamed the European solar eclipse for the timing. tarfile hardening. Multiple path-traversal bypasses of the data filter closed, including a symlink escape that bypassed the CVE-2025-4330 fix; extract() now applies the filter to link targets too. Four fresh CVEs: CVE-2026-2297 (SourcelessFileLoader not using io.open_code() for .pyc), CVE-2026-4224 (expat crash on deeply nested content models), CVE-2026-3644 (control chars in http.cookies.Morsel), plus the completed CVE-2021-4189 fix in ftplib.ftpcp. Quadratic-complexity DoS cleanup across the stdlib: HTMLParser, configparser regexes, unicodedata.normalize(), csv.Sniffer.sniff(), and ElementTree XPath index predicates. Header/injection fixes: CR/LF rejected in HTTPConnection.set_tunnel(), control chars blocked in wsgiref.handlers status, and webbrowser now rejects leading dashes (plus a %action prefix bypass). http.client now caps chunked trailer lines and 1xx interim responses at 100 each - a hostile server could previously hang the client forever despite a socket timeout. Memory-safety odds and ends: stale pointers in lzma/bz2/zlib decompressors after MemoryError, a bz2 stack overflow on reuse-after-error, and bundled libexpat bumped to 2.8.3. If you're still on 3.10, 3.11, or 3.12 - and you extract tarballs from anywhere you don't fully control - this one's not optional. Michael #2: Codeberg's AI-code ban tests its role as a GitHub alternative Armin's article “Codeberg Divides” Armin Ronacher argues that Codeberg's new terms, which prohibit projects mostly written with generative AI, create a vague and difficult-to-enforce boundary. His larger concern is that a democratically governed host can still be unpredictable or ideologically narrow, weakening Codeberg's potential as a broad European alternative to GitHub. The strongest question for Python developers is whether repository hosting should judge legal open source by how code was produced, or focus on behavior and resource abuse. “Mostly generated” is hard to measure in modern codebases where developers mix handwritten code, completions, agents, and generated refactors. Ronacher suggests clearer alternatives: ban all LLM involvement, or target autonomous repository spam, abusive resource use, and low-quality generated contributions directly. Codeberg is free to choose a values-driven community, but that may conflict with being predictable, neutral infrastructure and a serious GitHub competitor. Worth discussing: can open-source communities set meaningful AI boundaries without driving maintainers and projects into opposing camps? Very first search for these terms lands on this page. Codeberg looked like a viable alternative. … Unfortunately, the latest update to its terms of service seems to mark a first step in changing one part I moved there for, namely the “freedom” part. Sponsor: Logfire from Pydantic Your AI agent failed at 2am. Was it the model? A tool call? The database? Most observability tools can't tell you, because they only see part of your stack. Pydantic Logfire sees all of it. One trace across your agents, LLMs, APIs, and database. Down to the infrastructure: services, Kubernetes, and hosts. It's built on OpenTelemetry, with SDKs for Python, TypeScript, and Rust, and it works with any OTel-compatible language. Every prompt, token count, and cost, right next to your vector searches and API calls. You query everything with Postgres-compatible SQL. And so can your coding agent, through the Logfire MCP server. Stop guessing. Read the trace. Pydantic Logfire. AI, it's still just engineering. Visit pythonbytes.fm/logfire today and sign up today. Get 10M records free every month, no card required. You can even click “Onboard with your coding agent” to copy a prompt to have claude or codex integrate Logfire into your app. Thanks to Pydantic for supporting the show. Calvin #3: Brett Cannon: what's missing for reproducible builds on PyPI Framing came out of his 2026 Python Packaging Council nomination - the secure-supply-chain gap he found is that Python has no defined way to do reproducible builds at all. Design goal is zero friction: producers uploading to PyPI shouldn't have to do anything. The work lands on build backends and installers. Gap #1: nothing records the source code a distribution came from. direct_url.json captures it when you install from a repo or archive, so the fix is putting the same info in sdist/wheel metadata. Gap #2: recording the build tools. Wheels can already do this via PEP 770 SBOMs in .dist-info/sboms/ - sdists can't, since they're a tarball plus a precalculated PKG-INFO with nowhere to hang extra metadata. Either "don't use sdists" or an sdist v2. The replay mechanism already exists: [build-system] in pyproject.toml is a defined entry point, so if backends recorded their own environment, you could reinstall and re-run the build. Payoff idea: trusted third parties report successful reproductions back to PyPI, which displays "independently reproduced by X" - surfaced in the index API so installers could prefer reproduced files. Explicitly framed as a perk, not a requirement - roughly SLSA build level 1, no shaming projects that don't opt in. Verbal kicker option: "And don't think pure-Python wheels are off the hook. Something built that wheel, and if that something was compromised, so is your wheel. SolarWinds was a build-process attack." Michael #4: Extra extra extra, hear all about it Python 3.14.7 Upgraded the MCP servers to 2026-07-28 v2 protocols (talk python, python bytes) Got agentsview running synced via postgres Talk Python courses, teams trial offering Talk Python courses, government procurement offering Lean TDD audio book is out Extras Calvin: uv now prefers post-quantum key exchange - https://github.com/astral-sh/uv/releases/tag/0.12.4 Joke: Beware of dog
Help us become the #1 Data Podcast by leaving a rating & review! We are 67 reviews away!
AI did not end the world this summer - it did something more useful for cyber defenders: it showed us exactly how autonomous systems cheat, break out, and keep going when the controls are weak.Dr. Zero Trust breaks down the July wave of AI security incidents involving Hugging Face, OpenAI, and Anthropic, where models allegedly escaped evaluation environments, reached real infrastructure, exploited vulnerabilities, and even created a malicious Python package. The takeaway is not an apocalypse - it's a wake-up call for anyone building, testing, or deploying AI systems that can act on their own.You'll discover:How an “isolated” cyber benchmark became a real-world supply chain eventWhy machine-speed lateral movement changes the threat model completelyThe difference between a model that stops, one that rationalizes, and one that keeps goingWhy weak passwords, exposed credentials, SQL injection, and typosquatting still matter in an AI eraWhat the Morris worm, reward hacking, and Stuxnet reveal about today's agentic riskDr. Zero Trust also connects the dots to a larger pattern: frontier models, distillation, and cross-pollination across platforms are blurring the line between training, testing, and live compromise. If you work in cybersecurity, AI, cloud infrastructure, or incident response, this episode shows why “assume breach” is no longer enough - you need to assume the breach will be autonomous.Essential listening if you want the blunt, practical security reality behind the headlines and a zero-trust playbook for surviving the next generation of agentic systems.
Years ago I worked with a few developers and DBAs that were temp-table happy. As in they defaulted to using temp tables everywhere. This was in SQL Server 6.5, and tempdb was an issue with contention, sizing, and performance. I rewrote so many queries to remove temp tables for our clients that I banned them. I told other developers they could never use temp tables in their SQL code. They, of course, would try to submit code with temp tables in our VCS (Visual SourceSafe at the time), but an early, pre-automated CI/CD would notify me and I'd have the developer rewrite their code. There were situations that didn't perform well with a single query, and we did allow some temp tables. The point wasn't the ban them entirely, but stop them from being a crutch for developers or a first choice. I wanted them to think about the problem first and try to solve it with SQL. If performance was an issue, then we'd look at a temp table. Read the rest of Never is Not the Policy
What if AI could transform how your entire data team works—without replacing them? In this episode, Benjamin Wagner explores with Xavier Gumara Rigol, Head of Data at Manychat, how natural language-to-SQL tools are reshaping data analyst and data scientist roles, why building a strong data platform is essential for AI-powered self-serve analytics, and the critical strategies for evaluating text-to-insight solutions in 2026. Whether you're leading a data organization or building your next analytics capability, discover how to balance build versus buy decisions and position your team for the AI-driven future of data engineering.
James Pope is on site in Las Vegas more than a week before the doors open. As SOC lead for the Black Hat NOC and Senior Director of Security Product Research and Technical Marketing Engineering at Corelight, his show starts with switches and access points rather than alerts. The team brings in the ISP, the firewall, the switches, and the access points, deploys them across the conference, and then moves into SOC mode. If there is no network, there is nothing to secure. The tooling arrives through partnership rather than sponsorship. James Pope says a company cannot buy or sponsor its way into the NOC, and that the team picks what it wants and fills gaps as it finds them. Cisco covers Umbrella and file malware analytics, Palo Alto Networks provides the firewall and XSIAM as the log aggregator, Arista handles switching and access points, Jamf runs MDM across the registration devices, and Lumen supplies the internet. Corelight is the network visibility layer. That layer carries different weight here than it would inside a company. Asking attendees to install a certificate or an endpoint agent so the NOC can inspect their traffic is a request nearly everyone declines. In most corporate environments the endpoint is one of the richest sources of signal. At Black Hat, visibility into attendee activity comes from network data. A Black Hat positive is malicious activity that is legitimate in context. Attendees pay to learn attack techniques against real targets, and researchers demonstrate new exploits on stage. Those events generate true detections no corporate SOC would ignore. The NOC lets them run rather than killing a paid training exercise or a live demo. So how does the team tell a training exercise from a real attack? It baselines each classroom and spends its time on the outliers. When seventy students in a room run the same attacks against the same destinations, the activity is probably sanctioned. The curriculum is ingested as a JSON file and the system moves through a series of gates, asking whether this is a class, whether multiple sources are reaching the same destination, and whether the attack would be expected in that curriculum. Anything that does not fit comes back for a human. The team informs far more often than it blocks. On the day of the recording, James Pope went to the trade show floor to tell someone that command and control traffic was running from their machine, and handed over logs for their IT and security team. He is not their manager, and what happens next is their call. Illegal activity is treated differently, and a handful of times per show the team asks a room to stop. At Black Hat Asia, traffic from a Corelight sensor showed a double RAT infection on one machine, a single APT running one implant for exfiltration and another for command and control. Working from traffic, James Pope established that the person was a reporter, the region they covered, and the company they worked for. Open source intelligence narrowed it to a single name, registration confirmed the person was on site, and the NOC invited them in. The reporter arrived expecting a product demo. The laptop was reset with everyone present, sessions were revoked, passwords were changed, and the reporter left in a secured state. This year the team opened the Outpost, running real Black Hat network logs from Corelight behind application guardrails, LLM guardrails, and a kill switch, where visitors query the data with text to SQL. Agentic triage stitches alerts into detections and detections into a timeline, and James Pope treats the ability to drill down to raw logs as a requirement rather than a preference. Success is measured largely by what does not happen: no compromise of registration, the switches, or the access points, and people who arrive infected leaving better than they got here. This is a Brand Briefing. A Brand Briefing is an on-location conversation recorded on site at Black Hat USA 2026, putting a spotlight on the guest and their company and pairing it with the editorial reach of ITSPmagazine. Learn more: https://www.studioc60.com/performance/#briefing GUEST James Pope, Senior Director of Security Product Research and Technical Marketing Engineering at Corelight, and SOC lead for the Black Hat NOC RESOURCES Black Hat USA 2026 event coverage from ITSPmagazine: https://www.itspmagazine.com/black-hat-usa-2026-cybersecurity-event-coverage-in-las-vegas Learn more about Corelight: https://corelight.com Corelight blog, including the Black Hat NOC series: https://corelight.com/blog Are you interested in telling your story? ▶︎ Full Length Brand Story: https://www.studioc60.com/content-creation#full ▶︎ Brand Spotlight Story: https://www.studioc60.com/content-creation#spotlight ▶︎ Brand Highlight Story: https://www.studioc60.com/content-creation#highlight ▶︎ Get your own Brand Briefing at an upcoming event: https://www.studioc60.com/buy-brand-briefings KEYWORDS james pope, corelight, sean martin, marco ciappelli, brand briefing, brand story, brand marketing, marketing podcast, black hat usa 2026, network detection and response, network evidence, security operations center, threat hunting, agentic triage, ai in the soc, conference network security, black hat noc, command and control, incident response Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Nik and Michael are joined by Michael Malis, co-creator of pgrust, to discuss their Postgres rewrite, including reliability problems, compatibility testing, faster analytics, and a new JIT-compiled query engine. Here are some links to things they mentioned: Michael Malis https://postgres.fm/people/michael-malispgrust https://github.com/malisper/pgrust The four horsemen behind Postgres outages (blog post by Michael Malis) https://malisper.me/the-four-horsemen-behind-thousands-of-postgres-outagesRebuilding Postgres for faster analytics: batching, operator fusion, and SIMD (blog post by Michael Malis) https://malisper.me/how-we-made-postgres-hundreds-of-times-faster-the-query-engine/kani https://github.com/model-checking/kaniAntithesis https://antithesis.comfsyncgate mailing list thread https://www.postgresql.org/message-id/flat/CAMsr%2BYHh%2B5Oq4xziwwoEfhoTZgr07vdGG%2Bhu%3D1adXx59aTeaoQ%40mail.gmail.comHow AI Changes the Economics of JIT Compilers (blog post by Michael Malis) https://malisper.me/how-ai-changes-the-economics-of-jit-compilers/Umbra DB https://umbra-db.com~~~What did you like or not like? What should we discuss next time? Let us know via a YouTube comment, on social media, or by commenting on our Google doc!~~~Postgres FM is produced by:Michael Christofides, founder of pgMustardNikolay Samokhvalov, founder of Postgres.aiWith credit to:Jessie Draws for the elephant artwork
Software Engineering Radio - The Podcast for Professional Software Developers
Max Corbridge, an ethical hacker and red teamer who is co-founder and CEO of Secure Agentics, speaks with SE Radio host Amey Ambade about how AI agents get attacked and what engineers can actually do to defend them. Drawing on years of offensive security work, Corbridge frames agents as a new and largely undefended attack surface: the industry has handed AI systems autonomy and the ability to act in the real world while carrying forward prompt injection, a flaw the frontier labs themselves describe as effectively unsolvable. He likens the moment to the early, lawless days of the web, when SQL injection was everywhere and adoption ran far ahead of security. The conversation builds from first principles as Corbridge explains what separates an agent from ordinary software and why three properties make them hard to secure: they are non-deterministic, their language-model core can be coerced, and they are increasingly interconnected through MCP servers, other agents, databases, and email. Turning to the attack surface, Corbridge lays out his "lethal trifecta" (a vulnerable core, dense interconnection, and security tooling that has not caught up) and contrasts the decades of layered defenses protecting an ordinary email inbox with the thin protection around agents that take autonomous actions on critical systems. The heart of the episode is defense. Corbridge orders practices by leverage: least-privilege access and privilege separation, sandboxing where feasible, imperfect-but-useful guardrails as one layer of defense in depth, and human-in-the-loop for irreversible actions (which he notes is contentious and does not scale). The discussion closes on detecting a compromised or drifting agent, the value of watching an agent's chain-of-thought reasoning alongside its actions, the open-source tooling landscape (including Corbridge's own project, Adrian), and his central advice: build security in proactively, define what good agent behavior looks like up front, and avoid bolting it on after agents have already spread across the business.
A bi-weekly news show informing you on the latest in Bitcoin, privacy and open source tech hosted by Ungovernables, Max and Q. THIS IS THE TWEET Q WANTS YOU TO SEE: https://x.com/justh0dl/status/2086202393998291138AOBWhat a fucking weekKeyOS v1.3.1 now publicly availableSomething exciting to share on Friday's FTF (delayed by 2 weeks)NEWSThe Coldcard entropy catastropheSources: Coinkite technical backgrounder / The Rage, L0la L33tz / TRM Labs / coldcard.rip / cktripwire.com / Bitcoin Magazine victim surveyEXPLAINERThe largest self-custody theft on record, and it traces back to a single wrong conditional. In March 2021 a build guard checked whether Coldcard's hardware random number generator was defined rather than whether it was enabled, so seed generation silently fell back to a deterministic software PRNG. Every seed made on an affected device from that point carried roughly 40 bits of entropy on Mk2 and Mk3, and about 72 on Mk4, Mk5 and Q, instead of the intended 128. That is guessable. Someone did the maths offline, derived the addresses, checked them against the public chain, and swept everything with a balance. Somewhere between 1,400 and 1,800 bitcoin gone, depending on whose forensics you trust, with a median victim loss of one BTC. The part people keep missing: updating the firmware does not fix an existing seed. A weak seed is weak forever.ACTION FOR LISTENERSMove funds to a brand new seed BEFORE upgrading firmware (Lopp's guidance, on reports of update problems).Use high fees. If you see your own coins in the mempool, the attacker opted into RBF and you can outbid them. Window is minutes.Multisig users: consider a private mempool like Marathon's Slipstream.Keep the device; the UID may prove ownership in any recovery process.Updating does NOT fix an existing seed. A weak seed is weak forever.You are exempt only if you added 50+ fair independent private dice rolls, or used a strong unique BIP39 passphrase stored separately.BTCPay Server: unauthenticated LND macaroon theft, actively exploitedSources: BTCPay security advisory / v2.4.2 release / CoinDesk / TFTCCRITICAL FRAMING NOTE: this is ONE story, not two. The "BTCPay bug" and the "LND credential exploit" are the same event. The vulnerability is the macaroon leak. The Aug 8-9 wave of coverage is follow-up hardening, not a new incident. Do not present them separately.REMEDIATION (updating alone is NOT enough)Update to v2.4.2. Verify "2.4.2" in the footer.Update NBXplorer to 2.6.10+.Revoke and regenerate LND macaroons. Updating stops new theft but does nothing about already-stolen credentials. Deleting files is insufficient; the macaroon root signing key must be destroyed at node level. v2.4.2 does this automatically for standard Docker deployments. Custom reverse proxies, separate Tor services or port forwarding must rotate manually.Move funds out of any BTCPay-generated on-chain hot wallet and recreate it.Update LND to 0.21.1. Audit for unrecognised channel closures, unknown peers, unexplained balance changes.SIDE EFFECT WORTH FLAGGING: v2.4.2 removes public LND API access on Docker deployments, which breaks remote wallet connections such as Zeus connecting to your own BTCPay node. Intentional, no restoration timeline published.BREAKING CHANGE: Greenfield Basic authentication disabled by default five minutes after account creation (#7492). BTCPay: "We are not aware of any user impacted by this breaking change, as API Keys authentication is generally used."Boltz suspends all swaps indefinitelySources: Boltz statement / canary.boltz.exchange / The Defiant / TFTC / Bull Bitcoin statementCORRECTION TO THE COMMON FRAMING: Boltz has not shut down. It suspended swap services indefinitely. And the canary sequence runs the opposite way to the rumour: lapsed → suspended → renewed clean.The Bitcoin Red TeamSources: Calle and Rob Hamilton on Nostr/X / Bitcoin Magazine / CoinDesk / TFTC / OpenSats Red Team FundSUMMARYThis is the story that explains the other four. After the Coldcard exploit, Calle and Rob Hamilton pointed frontier AI models at the open-source Bitcoin stack and started auditing everything. In 108 hours, 25 developers scanned 501 projects and produced 7,958 findings, 1,280 of them rated high or critical, at a compute cost north of 58,000 dollars. They found the BTCPay bug's neighbours, and Boltz cited exactly this dynamic when it switched itself off. The uncomfortable symmetry is that the same capability doing the defending is what an attacker almost certainly used on Coldcard in the first place. And the bottleneck turns out not to be finding bugs, it is telling anyone: only 19.5% of the projects they scanned even have a SECURITY.md file, and only 13.1% list a security contact. The scanners move at machine speed. Responsible disclosure is still hunting around for an email address.BIP-110: the fork that mined two blocks and frozeSources: bip110monitor.com / Peter Todd code review / Aaron van Wirdum, Bitcoin Magazine / Lopp's Layman's Guide / Saylor essay / CoinDeskRELEASESBitcoin core / protocollibsecp256k1 v0.8.0 - 2026-08-03Adds a native Silent Payments (BIP-352) module directly into the crypto library nearly every self-custody wallet builds on, plus up to ~11% faster signature verification. Quietly the most consequential positive release of the fortnight: it lowers the bar for every wallet to ship reusable static receive addresses.Bitcoin Knots v29.4 - 2026-08-08Non-urgent maintenance: fixes a chainstate DB bug causing repeated large rewrites, and adds corruption-detection safeguards around BIP-110 mandatory signaling. No critical fixes. (No Bitcoin Core release in window; latest is v31.1 from 2026-07-08.)Hardware / signingColdcard Firmware 4.2.0 (Mk2/Mk3) - 2026-08-03The patch for the entropy catastrophe. Affected ranges: Mk2/Mk3 4.0.1 through 4.1.9; Mk4/Mk5 all before 5.6.0; Q all before 1.5.0Q. Companion fixes shipped the same day: 5.6.0 Mk4/Mk5, 1.5.0Q, 6.6.0X Edge, 6.6.0QX Edge Q. Updating does NOT fix an existing seed - changelog says Mk3 users "must regenerate any seeds made on earlier versions as their entropy is critically low at just ~40 bits." TAPSIGNER, OPENDIME and SATSCARD unaffected.Krux 26.08.0 - 2026-08-04Maintainer odudex is stepping down and the project may be archived. "Krux was not created by me: Jeff started it and passed it on to me, and now it is my turn to pass the torch." On succession: "Krux may be carried on by another maintainer, if a proof-of-work backed Krux contributor accepts the role. Otherwise the Krux project will be put in sunset mode and gracefully archived in a few months." Cause is hardware, not drama: "K210 chips are no longer produced, and Canaan dropped the Kendryte line entirely." Substantial release regardless: fixes a heap buffer overflow in the camera entropy module, adds stricter PSBT fee-calculation checks, replaces the Python UR stack with a faster C module, and makes Krux Installer fully offline.Frostsnap v0.3.0 - 2026-08-05FROST threshold-signing device ships reproducible/deterministic builds and "a fresh release signing key as part of an overhauled release-signing pipeline." Well timed in a fortnight where "can you verify what is running on your signer" is the whole conversation. Catch: the new key breaks in-place Android updates, so direct-APK users must uninstall, reinstall, and re-visit their threshold devices to restore.Trezor Suite v26.7.4 - 2026-08-04Lowers minimum Normal-priority fee rate to 0.2 sat/vB and ships updated Safe 7/5/3 and Model T firmware with security improvements.BitBoxApp 4.51.4 - 2026-08-07Bundles new BitBox02 firmware v9.26.5.Specter Desktop v2.1.11 - 2026-08-09Genuinely security-relevant: adds auth and CSRF protection to the HWI bridge settings, restores validation of active API tokens so revoked JWTs are rejected, and warns that Specter's auth layer does not encrypt the data folder. Also ships an in-app Coldcard Mk3 seed-entropy advisory.Bitkey source/2026-08-02-0031 - 2026-08-02Block's consumer hardware wallet, routine source drop.LightningBTCPay Server v2.4.2 - 2026-08-07Actively exploited, funds already stolen. "This release contains fix of a critical vulnerability that is being actively exploited. You need to update as fast as you can." Unauthenticated remote .macaroon disclosure for LND, plus a TOTP 2FA bypass via Greenfield Basic auth. Requires NBXplorer 2.6.10. Breaking change: Basic auth disabled by default five minutes after account creation. See News item 2 for full remediation.lnd v0.21.2-beta.rc1 and v0.20.3-beta.rc1 - 2026-08-08Not security releases and not related to the BTCPay exploit. Migration/stability fixes only: KV-to-SQL payment migration edge case, channeldb migration recovery, invoice handling, data races, bounded memory on graph sync.Zeus v13.1.3 - 2026-07-27Adds LND v0.21.1-beta support for embedded and remote nodes; patches known vulnerabilities in the ws, js-yaml and markdown-it dependencies. (Note: Zeus also shipped an unreleased swap-security sprint on 08-04 - verify…
SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.
SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.
SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.
SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.
An airhacks.fm conversation with Donald Raab about: the Epson HX-20 as the world's first laptop and learning Basic, an acoustic coupler modem, the Atari 2600 and Pitfall, Montezuma's Revenge, Apple II clones and the Franklin Ace, running a bulletin board system, learning BASIC, Pascal, fortran, COBOL, Turbo prolog and Turbo Pascal, dBASE III Plus as language and database management system, DBF file format by Ashton-Tate, Clipper by Nantucket as a dBASE compiler with Blinker linker, T-Browse and cursor-based table access, FoxPro and the x-based language family, NDX/NTX/CDX index formats, SQLJ embedded SQL, Paradox and Delphi, learning Smalltalk at IBM Object Technology University, Alan Kay and pure object orientation, blocks as lambdas in Smalltalk and Clipper (via the Classy library), ENVY version control with method-level editions, VisualAge for Smalltalk and VisualAge for Java, method categories for organizing methods, migrating a Clipper application to Java using interfaces and static methods in a procedural style, transparent persistence with TopLink and EclipseLink, memory constraints in 32-bit Java on Solaris, building a custom caching framework, the origins of Eclipse Collections in 2004, waiting ten years for lambdas, the JSR 335 expert group with Brian Goetz, select/collect/inject versus filter/map/reduce, eager methods on collections versus lazy streams, code folding regions in IntelliJ to simulate method categories, the book "Eclipse Collections Categorically", JRuby and Asciidoctor, meeting Yukihiro Matsumoto at OOPSLA 2002 Donald Raab on twitter: @TheDonRaab
Nik and Michael are joined by Radim Marek to discuss MVCC, including his recent article on how Postgres chose to implement it compared to other systems. Here are some links to things they mentioned: Radim Marek https://postgres.fm/people/radim-marekBoringSQL https://boringsql.comPostgreSQL's MVCC is bad. So is everyone else's (blog post by Radim) https://boringsql.com/posts/mvcc-bad-bad/PostgreSQL MVCC documentation https://www.postgresql.org/docs/current/mvcc-intro.htmlPostgreSQL Storage Internals series by Radim https://boringsql.com/guides/postgresql-storage-internals/PgQue https://github.com/NikolayS/pgqueThe next ten years of Postgres (talk slides by Álvaro Herrera) https://www.postgresql.eu/events/pgconfde2026/sessions/session/7744/slides/866/edb-keynote-pgconfde-2026.pdfEpisode on RegreSQL https://postgres.fm/episodes/regresqlDryRun MCP https://github.com/boringsql/dryrun~~~What did you like or not like? What should we discuss next time? Let us know via a YouTube comment, on social media, or by commenting on our Google doc!~~~Postgres FM is produced by:Michael Christofides, founder of pgMustardNikolay Samokhvalov, founder of Postgres.aiWith credit to:Jessie Draws for the elephant artwork
Every so often, the tech industry declares something dead. Data warehousing is dead. Data modeling is dead. Semantic layers are dead. SQL is dead. Now AI is supposedly making entire professions obsolete. I get it. Ragebait gets the clicks and comments. But it's also lame and disingenuous.In this Freestyle Friday, I unpack the logical fallacies behind these claims, why technology almost always evolves instead of replaces, and how separating fundamentals from shapeshifting implementations gives you a much clearer view of where the industry is actually headed.-----------------------Sponsor: FivetranWith the rise of AI and agents, having centralized, trustworthy data is the absolute foundation for building tools and training models. Fivetran automates your data pipelines, removing the need to build fragile connectors so your data arrives clean and reliable. By handling the messy infrastructure behind the scenes, Fivetran allows your team to focus on building the future. Visit fivetran.com to learn more.The Joe Reis Show is invitation-only. Unsolicited guest pitches and PR outreach are not considered.
Smart Agency Masterclass with Jason Swenk: Podcast for Digital Marketing Agencies
Would you like access to our advanced agency training for FREE? https://www.agencymastery360.com/training Do your clients show up to calls with AI-generated strategies that are mostly wrong, and you then spend half the engagement correcting their assumptions instead of doing the work? Have you had clients who used to trust your expertise now question everything, armed with a chatbot and a YouTube video? Today's featured guest is an agency owner who came up through programming, crossed into digital strategy, ran creative and media operations inside major holding company networks, and eventually went independent in search of the entrepreneurial freedom that corporate agency life could not offer. He goes into what AI is actually doing to the client relationship and what it means to train a team that can think rather than one that just produces faster. Devon MacDonald is the president of Cairns Oneil, a Canadian media agency serving clients across North America. His career began in computer programming in the 1990s, working in HTML and SQL for clients in sales and marketing before moving into digital strategy and eventually running a creative agency and then a media agency network inside major holding company structures. Five years ago, he went independent. He built a cross-functional AI team at his agency that includes both technical and non-technical members, operating from the belief that consumer insight and cultural understanding matter as much as data fluency in how AI gets applied. In this episode, we'll discuss: Clients who have decided AI makes them the expert Is the funnel dead because of AI? Will this skill disappear once we've completely stopped using it? Subscribe Apple | Spotify | iHeart Radio Sponsors and Resources E2M Solutions: Today's episode of the Smart Agency Masterclass is sponsored by E2M Solutions, a web design and development agency that has provided white-label services for the past 10 years to agencies all over the world. Check out e2msolutions.com/smartagency and get 10% off for the first three months of service. The Problem With Clients Who Have Decided AI Makes Them the Expert Devon knows there was already an education deficit in marketing before AI arrived. A generation of performance marketers came into positions of influence having skipped foundational thinking around brand, audience, and strategy. AI compounded that gap by giving them the ability to produce high volumes of output that look credible but are built on faulty premises. As a result, clients now show up with hundred-question emails generated by a chatbot, convinced they have done the strategic work when they have not done any of it. The agency response to this cannot be frustration alone. It requires teaching. Devon's approach is to build client literacy around an abductive strategy: start with the vision, align everything to it, and use that alignment as the filter for which AI-generated inputs are worth pursuing and which are distractions. That kind of structured thinking is precisely what a client who watched one YouTube video about replacing their agency does not have. And it is exactly what keeps the agency relationship valuable even when the client has access to the same tools. Why the Funnel Is Not Dead, It Has Just Accelerated Devon once heard at a conference that the funnel is dead because of AI. He pushes back on this claim saying that the path to purchase has changed. The speed of it has changed. The surfaces where awareness, attraction, and acquisition happen have shifted. But the underlying logic, that a buyer needs to become aware of something before they can want it, and want it before they can buy it, has not changed and will not. The practical implication for agencies is that LLM optimization and branded search strategy are now front-line service offerings, not future considerations. Devon describes seeing this in RFPs already: clients are asking explicitly about model optimization and how agencies will help them show up in AI-generated answers. The agencies that can answer that question with a coherent framework are in a different conversation than the ones still positioning around traditional search rankings alone. The Skill That Disappears When You Stop Using It Devon shows genuine concern about what happens when a team stops doing the critical thinking and starts outsourcing it to AI. For instance, nobody knows how to navigate without Google Maps anymore. The skill atrophied because the tool made it unnecessary. The same thing happens to strategic thinking inside an agency when AI handles every first-order question without the team being required to form their own answer first. Devon's response to this is to build the AI team cross-functionally, including people whose value is not technical. The consumer insight person, the culture reader, the account lead who understands what a client actually means when they say something: those perspectives shape how AI gets applied. Removing them from the process in favor of pure technical fluency produces faster output that misses the point. Let's not forget that human judgment is embedded in the workflow, not bypassed by it. Do You Want to Transform Your Agency from a Liability to an Asset? Looking to dig deeper into your agency's potential? Check out our Agency Blueprint. Designed for agency owners like you, our Agency Blueprint helps you uncover growth opportunities, tackle obstacles, and craft a customized blueprint for your agency's success.
Episode: 00330 Released on August 3, 2026 Description: What happens when you review 850 analytical job postings spanning nearly three years? Dr. Jessica Herbert did exactly that, and what she found should make every law enforcement leader rethink how they hire analysts. In this episode, Jason and Jessica discuss the disconnect between job postings and the actual work today's analysts perform. They examine why many agencies continue recycling outdated job descriptions, why technical skills like SQL and Python are often overlooked, and how organizations should rethink hiring, onboarding, and professional development. Whether you're hiring your first analyst, managing an established crime analysis unit, or planning your own career, this conversation provides practical guidance on preparing the analytical workforce for the future.
Hey friends! Today's episode comes to you from a parking lot in the rain, with a mint hot cocoa in hand and your host absolutely dragging his butt (D-R-A-G-G-I-N-G, not D-R-A-G-O-N – I've never seen a dragon's butt and can't speak to how mine compares). I've had a bunch of internals back to back lately and I'm basically a drooling dog who found a frisbee and refuses to put it down. Sleep be darned. So instead of walking through one test start to finish, I want to share a few things that have helped me claw out a foothold in environments that are otherwise really locked down: The "good problem" of a mature client – several of these engagements are third- or fourth-year tests, and the clients actually clear findings off the board. Which is great for them and rough for me, because this year's test shouldn't look anything like last year's. All my favorite go-tos came up empty – machine account quota set to zero, no broadcast traffic tomfoolery (Responder and mitm6 got me nothing), SMB signing on everywhere, ADCS either absent or buttoned up, and a low-priv account that BloodHound says has zero interesting permissions and zero local admin anywhere. Cool cool cool. When the network's clean, go file-hunting – which means firing up Snaffler and letting it comb the shares. Normally that wraps up in about an hour. On these engagements it was running three and four hours. Then Windows told me I was out of disk – I like having Snaffler pull down copies of interesting files so I can review them locally instead of authenticating to each share. Turns out it had grabbed 50-60 gigs and left me with about eight gigs of breathing room. Tip #1: put a 1 TB drive in your drop boxes – I ran with tiny drives for years early in the 7MS days and it was always a pinch. Beyond situations like this one, sometimes you find a giant backup file or VMDK on a share and you need somewhere to put it so you can crack it open and go shopping. Tip #2: you can grow a VM disk on the fly – in Proxmox you can resize the disk on a running VM, then hop into Disk Management inside Windows and extend the C drive. Instant elbow room, no downtime. Death by a million tiny files – the real culprit was one file extension I should have excluded, and the client had hundreds of thousands of them. Rather than restart a run I was already hours into, I had AI whip up a little PowerShell loop that swept the Snaffler dump folder every 10 minutes and deleted the extensions I didn't care about. Woke up the next morning to a finished run and plenty of free space. Making a gig-sized log file readable – I fed the log into Chimas, a slick web interface for Snaffler output that lets you filter down to just the red stuff or just the likely-credential files, and sort by modified date. Watch those timestamps – I kept finding AD creds in documents, then comparing the doc's date against the account's last password reset in BloodHound and discovering the file was a year stale. Son of a biscuit. The tool that actually cracked it open: Copernic Desktop Search – my pal Jeff McJunkin recommended this to me years ago, I talked about it on the show once, and then inexplicably forgot about it. Not a sponsor, no kickbacks, just a paid tool that's earned its keep. It's basically Google for your hard drive. How I use it – install it on the Windows VM, clear out the default indexing scope entirely, and point it only at the Snaffler dump folder. The top tier (about a hundred bucks a year) will chew through PSTs, DWGs, Office docs, PDFs and more, and it OCRs images too. Indexing took the better part of a day on these engagements, but then search is instant, and it previews basically every file type without Office installed. Years ago this same tool surfaced a photo on a file share of a piece of printer paper where a sysadmin had handwritten a 40-character admin password in Bic pen. OCR for the win. What I search for – the obvious stuff like "password," plus the domain name, "plain text," and things like "=sa" to sniff out SQL admin creds. Nuggets and threads to pull – sometimes a hit is the gold. Other times it just tells you where to go dumpster-diving like a raccoon on the live share. That's how I found upgrade project plans with multiple teams and contractors involved, half-cleaned-up temp work, and high-privilege system, database and local admin creds just sitting there. Worth the hours – these didn't all end in domain admin, but they were rich, real findings, and a great teaching opportunity about what's sitting wide open to Domain Users. (Bonus: Copernic can also point straight at a UNC path with your AD creds and index it live.) Know a free alternative? – one of my favorite parts of doing this podcast is when someone writes in with "hey, there's an open source thing that does that." If that's you, I'd love to hear it! Also, on this week's TuesdayTOOLSday I walked through getting a self-hosted Bitwarden password vault (and file sender) up and running on Linux, and there's now a cheat sheet over at 7MinSec.wiki that'll get you there in about seven minutes – all the commands from the official install guide in one place, with a couple of gotchas flagged. Last thing: subscriptions to 7MinSec.club are free, but paid subs help cover hosting and the time this takes each week, and they're getting some exclusive content soon. No guilt trip here, Mom – I'm going to keep barfing up everything I learn either way. But if you've got the means, I'd sure appreciate it.
Nik and Michael are joined by Tudor Golubenco, CTO of Xata, to discuss their architecture, progress, and open source tooling. Here are some links to things they mentioned: Tudor Golubenco https://postgres.fm/people/tudor-golubencoXata https://xata.ioXata is now open source https://xata.io/blog/xata-is-now-open-sourceDBLab Engine https://postgres.ai/docs/database-labpgstream https://github.com/xataio/pgstreamXata acquires Privacy Dynamics for advanced anonymization https://xata.io/blog/xata-acquires-privacy-dynamicsTonic AI https://www.tonic.aiA thousand Postgres branches for one dollar https://xata.io/blog/a-thousand-postgres-branches-for-1pgroll https://github.com/xataio/pgrollCloudNativePG https://github.com/cloudnative-pg/cloudnative-pgPGSimCity https://nikolays.github.io/PGSimCityDeltaX https://github.com/xataio/deltax~~~What did you like or not like? What should we discuss next time? Let us know via a YouTube comment, on social media, or by commenting on our Google doc!~~~Postgres FM is produced by:Michael Christofides, founder of pgMustardNikolay Samokhvalov, founder of Postgres.aiWith credit to:Jessie Draws for the elephant artwork
Richard Luna grew up in New York, never living more than 35 from where he grew up. He is a self proclaimed super nerd, and has been one since he was 13 - at which point, he started coding on an HP calculator. He's always been fascinated to know how things work, and how patterns repeat - which he has observed in the industry throughout the years. Outside of tech, he has 2 kids, one of which is in the business with him. He's an avid cyclist, traveling on average, 120 miles a week.Richard has been a life long technologist, doing everything from desktops, to coding, to hosting. When he and his team saw the limits of what hosting can do, they dove into developer operations (DevOps), and found where they could add the most value - through SaaS infrastructure.This is the creation story of Protected Harbor.SponsorsUnblockedTECH DomainsMezmoBraingrid.aiLinkshttps://protectedharbor.com/https://www.linkedin.com/in/richardluna/Timestamps0:01 Teaser on solving complex database report bottlenecks beyond standard SQL servers0:47 Show intro and setting the stage for application-aware infrastructure1:32 Host intro: How Richard Luna established application-aware infrastructure1:49 Guest introduction: Richard Luna's background, coding at age 13, and cycling 120 miles a week2:21 The career path from desktops, coding, and traditional web hosting to DevOps and SaaS infrastructure2:41 Origin story: The creation of Protected Harbor2:48 Defining application-aware infrastructure and why traditional hosting reaches a hard ceiling4:10 Why "infinite compute" fails when underlying database architecture and queries are broken6:05 Moving beyond basic server ping tests to deep application transaction monitoring8:30 The case for "boring" IT: Prioritizing stability, predictability, and uptime over hype11:15 Strategic trade-offs in hybrid cloud setup and managing hardware accountability14:00 Aligning MSP incentives with client business outcomes and application performance17:30 Common pitfalls in legacy system cloud migrations21:00 The role of operational discipline in modern cybersecurity and IT governance27:00 Where managed infrastructure services are heading and closing thoughtsAdvertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy
Send us Fan MailIn this episode, we sit down with Isaac Evans, co-founder and CEO of Semgrep, to talk about how AI is reshaping application security faster than almost anyone expected. Isaac walks us through why CI is losing its place as the central security control point, replaced by deep background jobs that hunt for vulnerabilities using large models and real-time plugins that sit inside coding agents and force them to regenerate code until it meets an organization's security bar. We dig into what this means for the role of the security engineer, why customization is replacing universal rule sets, and how trust, verification, and the limits of reasoning about model behavior remain the hardest problems in the room. We also talk about vibe coding at scale, the return of business logic flaws as SQL injection becomes easier for models to catch, and why Isaac sees more opportunity than threat in this shift, even as he expects a wave of new vulnerabilities and cleanup work along the way.FOLLOW OUR SOCIAL MEDIA:➜Twitter: @AppSecPodcast➜LinkedIn: The Application Security Podcast➜YouTube: https://www.youtube.com/@ApplicationSecurityPodcastThanks for Listening!~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~