POPULARITY
Categories
Veo scooters,bikes, and trikes are ready to ride on nearly every corner in Denver, but whose job is it to enforce their rules of the road? Contributor Michelle Jackson joins guest host Adrian Felix to discuss who should be accountable for scooters on the sidewalk. Then, a new program is making the ballet more accessible for everyone. Could other entertainment venues and organizations follow suit? A special members only segment dives into Denverites potentially being burned out on the mayor race already; become a City Cast Denver Neighbor today to listen! Plus, we end the week with a classic round of wins and fails. Michelle brought up the low water levels at Dillon Reservoir. Adrian discussed Marshall Zelinger's conversation with Congressman Gabe Evans, and a new champagne bar coming soon. For even more news from around the city, subscribe to our morning newsletter at denver.citycast.fm. Follow us on Instagram: @citycastdenver Chat with other listeners on Reddit: r/CityCastDenver Support City Cast Denver by becoming a member! What show would you like to see for cheap? Text or leave us a voicemail with your name and neighborhood, and you might hear it on the show: 720-500-5418 Learn more about the sponsors of this August 14th episode: Chalat Hatten and Banker Blue Sky Denver Health Regional Air Quality Council Looking to advertise on City Cast Denver? Check out our options for podcast and newsletter ads at citycast.fm/advertise
BMW spielt in Kooperation mit Sony Pictures in über 70 Ländern einen Animationsclip zum Spider-Man-Filmstart auf dem Bildschirm im Auto aus, freigeschaltet per Banner-Klick in bestimmten Modellen. BMW nennt es eine besondere Überraschung und ausdrücklich keine Werbung, in Foren beschweren sich Fahrerinnen und Fahrer trotzdem. BMW-Manager Durach hatte das Auto 2023 noch als letzten privaten Rückzugsort bezeichnet. Ich ordne ein, was das für Werbung im Auto und für das Vertrauen der eigenen Community bedeutet.Außerdem in den Marken- und Marketingnews der KW 33:⚽ DFB-Pokal Rebranding 2026: Kurz vor der ersten Hauptrunde am 21. August 2026 baut der DFB die Pokalmarke mit Strichpunkt und Manera um. Der Wettbewerb wird künftig über die gesamte Saison erzählt statt nur über den Finaltag, digital in kräftigem Violett, in den Stadien weiter in Pokalgrün.
Tras la subida de los seguidores en redes sociales en tiempo record y teneros en cuenta he decidido que voy a compartir con vosotras el hábito maestro. No sabía que correr fuera tan aspiracional, y sí, nunca es tarde para empezar. Veo muchos errores a la hora de adquirir este maravilloso hábito y aquí estoy para enseñarte cómo hacerlo bien. Voy a lanzar un programa para principiantes, para las que no sabéis por donde empezar y sobre todo para que os enamoréis de este hábito. Apúntate a la newsltter gratuita para saber más: www.amagoiaiezaguirre.es
Don, Aaron and Jack unpack a big week in venture, AI and building companies. Venture Downunder with Innovation Bay: pirate parties, poker, the Kill All The Puppies triage session (bridge, dilute, sell or fold), and why founder pitches at events should be three minutes, not ten. Canva blames frontier AI costs for a rare revenue downgrade: growth guidance cut from 30% to 20%, Canva AI 2.0 halted, in-house models to replace OpenAI and Veo 3, and why Tribe cancelled its Canva seats for Claude Design. Sovereign AI: Prof Anton Van Den Hengel's case for Australian capability. The only capability that survives is research, your advantage is momentum not checkpoints, and Anthropic engineers reckon they have nine to twelve months of coding left. HotDoc founder Ben Hurst steps away after 14 years: the maturity to recognise when your company needs a CEO and not a founder, and why the two energies are different jobs. Incumbents milking it: Don's Ansarada rant (pricing that scales with your data room, a 2007 interface, switching costs locked in by your deal risk framework), and Airtable reportedly selling for about $1 billion, 88% below peak, with founders wiped out by the preference stack. How 26-year-old Grace Brown built Andromeda into a $100 million robotics company, and the story of Abby the robot getting a couple dancing to their wedding song again after 20 years. Anti-portfolio goals: Vu Tran on why making a top VC's anti-portfolio is a benchmark, not a rejection. Ken Griffin on Citadel buying Sowood's $30 billion portfolio overnight: things come to those who wait, but only what's left by those who hustle. Sharts: data centre investment overtaking hotels and retail, Peter V'landys and NRL scoring, Federal Circuit Court general protections claims up 100% in three years, and EVs plus hybrids about to cross petrol and diesel in Australian vehicle sales. hello@tribeglobal.vc
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
This week on Electrek's Wheel-E podcast, we discuss the most popular news stories from the world of electric bikes and other nontraditional electric vehicles. This time, that includes new e-bikes from VW and Specialized, a Huffy kids' e-bike, Veo adds a three-wheeler to its shared micromobility lineup, we tested the new Juiced Scrambler, Anyland Rev Pro HMP launches 75 mph Flash e-scooter, and more.
Gálatas 4:29 "Pero como entonces el que había nacido según la carne perseguía al que había nacido según el Espíritu, así también ahora. " Quiero invitarlos a analizar la historia y meternos en la intimidad de una de las familias más conocidas de la Biblia: La familia de Abraham y Sara:Dios les había dado una gran promesa, pero los años pasaban y nada que se cumplía. Así que, en medio del desespero, decidieron hacer lo que muchos de nosotros hacemos hoy en día: intentar "ayudarle" a Dios con una solución puramente humana.De esa decisión nace Ismael. Fue un hijo producto de la carne, del afán y de la lógica humana, no del diseño del Espíritu. Sara y Abraham armaron una solución que, por encima, parecía resolver el problema, pero que en el fondo era frágil y trajo mucho dolor tanto en la vida de Sara como también en la vida a Agar. Construyeron una solución "de cartón". Y las consecuencias de intentar forzar las cosas en sus propias fuerzas, en lugar de esperar la promesa de Dios, generaron conflictos familiares tan profundos que aún en el día de hoy podemos ver las consecuencias.Precisamente sobre este contraste entre lo que hacemos en nuestras fuerzas y lo que Dios hace a través de su Espíritu, es que el apóstol Pablo nos habla hoy en su carta a los Gálatas: "Pero como entonces el que había nacido según la carne perseguía al que había nacido según el Espíritu, así también ahora." — Gálatas 4:29 (RVR1960)Es interesante ver los dos planos que el apóstol Pablo propone para llevar a los gálatas a reflexionar; los lleva a analizar cuándo algo proviene del Espíritu y cuándo algo proviene de la carne.Por eso, me centré en este pasaje para estudiar un poco a esta familia, ya que el apóstol Pablo la trae a colación. Vamos a revisar los peligros que amenazan a las familias, esos desafíos por los cuales todos atravesamos en estos tiempos.A menudo nos enseñan cómo construir una familia, cómo formarla y cómo convivir en ella, pero solemos perder de vista los errores que nosotros mismos generamos. A veces queremos que nuestras familias mejoren y sean transformadas, pero los cambios que proponemos son meramente humanos; y cuando son de origen humano, lastimosamente no son permanentes, sino pasajeros.Lo explico de esta manera: si intentas imponer reglas en tu hogar motivado por la ira, el resentimiento o el mal carácter, tal vez consigas resultados temporales, pero no lograrás un cambio real. Podrás infundir temor, pero nunca una verdadera transformación. Y cuando escudriñamos la Palabra para entender por qué sucede esto, descubrimos una gran verdad que hoy quiero compartir contigo y con todas las familias que me escuchan: la única forma de que un hogar sea verdaderamente transformado es a través del Espíritu Santo.Podemos acumular riquezas, alcanzar un alto nivel educativo, tener prestigio social o mucho reconocimiento, pero hoy vengo a decirte que la única manera de edificar una familia sólida es mediante el poder del Espíritu Santo.Si deseas tener una buena familia, necesitas a Cristo. Todo lo demás es bienvenido (los recursos económicos, que tus hijos estudien en un excelente colegio o en una gran universidad), pero absolutamente nada podrá reemplazar la presencia de Dios en un hogar. Cuando permites que la presencia de Dios reine en tu casa, tendrás una familia bendecida para Su gloria. Si Cristo está en la casa, la familia avanza. Cuando Cristo gobierna el hogar, este trasciende hacia el propósito que Dios ha trazado para él. A la manera de Dios o a mi maneraAhora, ¿qué pasa cuando los cambios y la transformación en una familia no nacen del Espíritu? Entonces provienen de tus opiniones, provienen de tu crianza. Muchas veces decimos: "Es que así me criaron a mí". Entonces, como tu papá cogía un palo, tú ya también tienes un palo; y como tu mamá tenía una chancla, tú también andas por ahí armado con una chancla en la cintura porque así te criaron. Hoy no vine a hablarte de tu crianza, hoy vine a hablarte del poder del Espíritu de Dios para que transforme a la familia. No vine a hablarte de lo que viviste en los tiempos de tus abuelos, vine a decirte que Dios está interesado en tu familia para traer una transformación a través del Espíritu, y no por métodos humanos.Cuando permites que Cristo intervenga en tu familia, entonces ya no serán tus opiniones ni tus emociones las que manden, porque actuar bajo nuestras emociones hace que, en lugar de arreglar las cosas, empeoren mucho más. Lo que nace de la carne se daña, pero lo que es del Espíritu permanece. Por eso te decía que esos cambios que hacemos en nuestra carne, si son normas, están muy bien, pero hay que ungirlas con el Espíritu Santo en oración, hay que pedir que Dios nos guíe y nos hable para que las cosas realmente puedan mejorar, porque cuando lo haces solo a tu manera, el hogar va por un camino en cierta temporada y por otro muy distinto en otra.Veo en la Biblia que el apóstol Pablo menciona a Isaac y a Ismael para hacer el paralelo entre cuando Dios mete la mano y cuando el hombre mete la mano. Si hemos traído al psicólogo para que nos ayude en casa, no está mal. Si hemos acudido al médico, no está mal; o al arquitecto... Hacemos tantas cosas por nuestra casa en lo físico, en la estructura, en lo mental y en lo emocional. Pero hoy vine a decirte: invita a Cristo a que viva en tu casa, no solo a que la visite. Cuando Cristo vive en la casa, a la casa le va bien, es bendecida, los hijos son bendecidos, el matrimonio es bendecido. No a tu manera, sino a la manera de Dios.Como un ejercicio pedagógico y para que lo recuerdes, quiero que lo pienses así: vamos a definir si en nuestra casa estamos haciendo las cosas a la manera de Dios o a tu manera. Tenemos la tendencia a querer hacerlo a nuestra manera, nos gusta, entonces nos imponemos y decimos cosas como: "¡Aquí se va a hacer así!". Ese tipo de actitudes demuestran que en esa casa se están haciendo las cosas a tu manera. Y hoy te estoy hablando de hacerlas a la manera de Dios. El Gobierno del Espíritu en el HogarCuando estaba estudiando un poco la historia de Isaac e Ismael, yo decía: "Señor, ¿cómo es posible que estos dos nunca pudieron estar en paz?". Ismael venía de un actuar de la carne, y ese actuar siempre terminó persiguiendo a lo que venía del Espíritu. Cuando en una casa no se le entrega el gobierno a Dios para que el Espíritu de Dios la gobierne, entonces otras cosas comienzan a gobernar: no gobierna el Espíritu de Dios, gobierna tu temperamento; no gobierna el Espíritu de Dios, gobiernan tus ideas o lo que tú quieres.Hoy vengo a decirte que, cuando el Espíritu Santo está en tu vida, lo que va a gobernar no son tus deseos, sino los frutos del Espíritu Santo. Podrás tener amor por tus hijos, habrá gozo y paz en el entorno. Ve, estudia los frutos del Espíritu, y te darás cuenta de lo que es una verdadera familia y una verdadera casa. Porque cuando una familia se construye sin Dios, termina luchando contra lo que viene de Dios.A una familia que no está construida por el Señor no le gusta orar, porque eso viene del Espíritu. No le gusta hacer el devocional, porque eso viene del Espíritu. No le gustan las normas ni hacer lo correcto, porque eso viene del Espíritu. Entonces los hijos quieren entrar a la hora que les da la gana a la casa, porque eso no viene de Él, viene de cómo se está manejando nuestra casa según la carne. "Es que todos los amiguitos en el colegio...", para ahí. "Es que en la universidad se hace...", pero es que nuestra casa no puede estar ...
Veo una entrevista a Milton Friedman, grabada en el año de 1979. Tiene una frescura y una actualidad extraordinarias.
In this episode, I sit down with VEO (@vrexec) — American expat living in Europe, self-employed consultant, investor, advisor, father, husband, and relentless optimizer. We dive deep into what it really takes to thrive as a location-independent professional: leaving high-stress corporate paths, getting your body and mind in peak shape, navigating life in Europe, and building a resilient self-employed lifestyle that actually works. From family construction roots to deal-making across borders, cultural contrasts between America and the Old World, daily systems for staying informed without doom-scrolling, and practical lessons on entrepreneurship and fatherhood while traveling — this is a no-fluff conversation for digital nomads, expats, flag theory practitioners, and anyone considering the leap to a more mobile, sovereign life. Whether you're in Latin America, Europe, or still planning your exit, VEO shares hard-earned wisdom on peak performance, opportunity hunting, and why now is the best time to get your life in order.
¿Por qué Moisés invierte todo un capítulo para hablar de la descendencia de Esaú? En hermenéutica está el principio que “a más información mayor importancia”. ¿Qué tendría de importante la descendencia de Esaú que Moisés se toma el tiempo para invertir todo un capítulo en la descripción de ella?Veo en este capítulo algunas lecciones importantes que aprender.
Veo cómo cada vez más personas convierten el deporte en una fuente de estrés, enojo y discusiones. En este episodio reflexiono sobre una escena que nos recuerda el verdadero propósito de competir: disfrutar, crecer y compartir. Porque incluso los atletas de alto rendimiento entienden que relajarse también forma parte del éxito. La pasión por un equipo o un deporte nunca debería costarte la tranquilidad. Dale a la
Veo el Mañana Recordando el Pasado por Bishop Joaquin G. Molina
Goodbye, FirstBank! While PNC promised a smooth transition after it acquired the legacy “Colorado bank for you” for $4.1 billion, many customers are reporting real problems – like the woman who suddenly had access to her mother's bank accounts. Green chile correspondent Justine Sandoval is back on the pod with host Bree Davies to talk about local banking brand loyalty, long wait times at the DMV, and the city's plan to staff up and write more parking tickets. Plus, a bonus segment only for City Cast Denver Neighbors about why DPS superintendent Alex Marrero can't seem to keep his desire to get out of Denver a secret. Join the Neighborhood to get all the goss and support this show by signing up to become a member today! Bree discussed our episode on population trends, Joy's Kitchen's new location, and the child who died in a Veo e-bike accident last month. Justine mentioned Colorado's struggle to find geriatric care doctors and Kelsey Pfendler's record-breaking solo boat row from California to Hawaii. Check out City Cast's national show, Your City Could Be Better! For even more news from around the city, subscribe to our morning newsletter at denver.citycast.fm. Follow us on Instagram: @citycastdenver Chat with other listeners on reddit: r/CityCastDenver Support City Cast Denver by becoming a member! Do you have a nightmarish – or pleasant – DMV story? Text or leave us a voicemail with your name and neighborhood, and you might hear it on the show: 720-500-5418 If you enjoyed this interview with Danny Feely, the Director of FP&A at TaskRabbit, learn more here. Learn more about the sponsors of this July 10th episode: Denver Art Museum Blue Sky Regional Air Quality Council Denver Botanic Gardens Energy Outreach Colorado Looking to advertise on City Cast Denver? Check out our options for podcast and newsletter ads at citycast.fm/advertise
Join our mastermind community: https://www.skool.com/apparel-success-mastermindTry the best Ai design platform: https://www.design.com/rob88AI video generators are getting insanely realistic, and in this video I show how clothing brand owners can use tools like Kling AI, Seedance, Google Veo, Sora, Runway, Adobe Firefly, Poyo.ai, Higgsfield AI, ChatGPT, ElevenLabs, CapCut, and Premiere Pro to create realistic ads, TikTok videos, Instagram Reels, product videos, and social media content without spending thousands on photoshoots, models, or videographers.I tested the biggest AI video tools to see which ones worked best for clothing brands using real product reference images. I break down why Google Veo, Sora, and Runway struggled, why Kling AI created some of the most realistic results, and why Seedance 2.0 might be one of the best AI video generators for accurate clothing brand content, fabric, logos, product details, and lifestyle scenes.If you run a streetwear brand, gymwear brand, hoodie brand, outdoor brand, or apparel business, this video shows how AI can help you create better ads, Meta ads, TikTok ads, Instagram content, UGC-style videos, and product marketing faster than ever.
ChatGPT tasks are back, Jack. ✅While we were collectively ping-ponging the Anthropic vs. U.S. government saga, the big tech AI players rolled out a TON of fresh AI features that are available today. ↳ Claude Design got a big upgrade↳ Google Vids got some serious AI sparkle↳ And there's a new Open Weights model king We'll break it all down. ChatGPT's Task Comeback, Claude's Design upgrade, Codex Copies your workflow and 7 other Fresh AI features you'll Want to use Today -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:ChatGPT Pulse Sunsetting and Tasks Comeback2. ChatGPT Scheduled Tasks Features and Access Tiers3. Claude Design June Update Overview4. WYSIWYG Editing and Design System Imports in Claude Design5. Claude Design Export Options and Third-Party Integrations6. Google Vids AI Avatars Upgrade with Veo 3.17. OpenRouter Fusion Multi-Model Synthesis Feature8. Claude Code Artifacts for Team and Enterprise Plans9. GLM 5.2 from ZAI Open Weights Model Overview10. GLM 5.2 Benchmarks and Enterprise Use Cases11. OpenAI Codex Record and Replay Feature Explained12. Codex Record and Replay vs. Traditional RPA ToolsTimestamps:00:00 Intro: 7 new AI features you can use today02:35 ChatGPT Tasks: Pulse is gone, Tasks are back04:31 Who has access to ChatGPT Tasks08:18 Claude Design June update overview09:14 WYSIWYG editing and Claude Code integration12:07 Claude Design export options and third-party integrations15:21 Google Vids AI Avatars upgrade18:15 OpenRouter Fusion multi-model synthesis21:54 Claude Code Artifacts for teams25:36 GLM 5.2 from ZAI open weights model29:02 OpenAI Codex Record and ReplayKeywords: ChatGPT Tasks, ChatGPT Pulse, OpenAI, scheduled tasks, proactive AI agent, Claude Design, WYSIWYG editor, Claude Code, design system import, PowerPoint export, Google Vids, AI avatars, Veo 3.1, Gemini 3.1 Flash, OpenRouter Fusion, model fusion, multi-model synthesis, Claude Code Artifacts, Claude Team plan, GLM 5.2, ZAI, open weights, MIT license, mixture of experts, Codex Record and Replay, RPA, workflow automation, Artificial Analysis, Hugging Face, Canva integrationSend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist.
CADENA 100 informa de la actualidad. Se esperan tormentas en el norte, interior y Aragón, con altas temperaturas. Begoña Gómez declara hoy en el juzgado por corrupción y tráfico de influencias; Zapatero es citado el miércoles por el caso Plus Ultra. Reino Unido prohíbe redes sociales a menores de 18, salvo WhatsApp. Se celebra el centenario de Marilyn Monroe con un récord de disfraces. Se descubre que antepasados usaban el fuego un millón de años antes. Pablo Gallinar cumple su sueño de vuelo acrobático. '¡Buenos días, Javi y Mar!' lanza "Misión Posible", invitando a oyentes a compartir sueños inalcanzables con dinero, como ser actriz o abrazar un chimpancé. Sanidad amplía el plan Veo de ayudas para gafas y lentillas a menores de 16. Un triatlón en Valencia se interrumpe por medusas. Oyentes comparten anécdotas de sus hijos con móviles, desde llamadas internacionales a suscripciones. De Pol presenta nuevo tema, Gente de Zona anuncia gira y Dani Fernández regresa a los ...
Africa is the literal center of the world's map and increasingly the center of gravity for ISIS, the manpower source for Russia's war in Ukraine, and the contested geopolitical ground where China builds bases and drops off free weapons. Our first active-duty guest pulls back the curtain on a combatant command that runs on 0.1% of the defense budget. LTG John W. Brennan Jr. is Deputy Commander of U.S. Africa Command and a 30-year career Special Forces officer, with command tours spanning 5th Special Forces Group, the anti-ISIS task force in Syria, and 1st Special Forces Command. He's joined by ChinaTalk's Justin, who served under Brennan as a young NCO in the Middle East. We discuss… How AFRICOM runs a counter-VEO away game on 0.1% of the defense budget by working “by, with, and through” partners “Putin's Purse”: trafficking thousands of Africans onto the Ukrainian front lines under false pretenses The Houthi–al-Shabaab pipeline and the threat triangle around Djibouti's PRC naval base Building an “alternate DIB in exile”: drone centers of excellence in Morocco, South African artillery, Namibian satellite radios Why Brennan wants to “declare jihad against proprietary data streams” and where AI actually helps a combatant commander decide WarTalk's first Ivorian dance party suno song: https://suno.com/s/1hhJTtwBn2NGR8eT Learn more about your ad choices. Visit megaphone.fm/adchoices
Africa is the literal center of the world's map and increasingly the center of gravity for ISIS, the manpower source for Russia's war in Ukraine, and the contested geopolitical ground where China builds bases and drops off free weapons. Our first active-duty guest pulls back the curtain on a combatant command that runs on 0.1% of the defense budget. LTG John W. Brennan Jr. is Deputy Commander of U.S. Africa Command and a 30-year career Special Forces officer, with command tours spanning 5th Special Forces Group, the anti-ISIS task force in Syria, and 1st Special Forces Command. He's joined by ChinaTalk's Justin, who served under Brennan as a young NCO in the Middle East. We discuss… How AFRICOM runs a counter-VEO away game on 0.1% of the defense budget by working “by, with, and through” partners “Putin's Purse”: trafficking thousands of Africans onto the Ukrainian front lines under false pretenses The Houthi–al-Shabaab pipeline and the threat triangle around Djibouti's PRC naval base Building an “alternate DIB in exile”: drone centers of excellence in Morocco, South African artillery, Namibian satellite radios Why Brennan wants to “declare jihad against proprietary data streams” and where AI actually helps a combatant commander decide WarTalk's first Ivorian dance party suno song: https://suno.com/s/1hhJTtwBn2NGR8eT Learn more about your ad choices. Visit megaphone.fm/adchoices
I don't know about you, but to me there are few things as interesting as the hardware/software interface: the point where carefully written code meets the messy, physical world of sensors, lenses, and real-time constraints. It's where a clever abstraction either holds up or falls apart the moment a real signal hits it.That makes Veo a perfect guest. The Copenhagen-based company builds AI-powered cameras that record and analyze sports matches, from grassroots football pitches to professional clubs, and then turn hours of raw footage into something coaches and players can actually use: automatic highlights, player tracking, and match analysis. To get there, they have to capture panoramic video on a custom camera, follow the action without an operator, and crunch an enormous amount of data, reliably and at scale.My guests sit on both sides of that interface. Anders Hellerup Madsen works close to the metal on the camera itself, on the embedded firmware and the GStreamer media pipeline that turns raw sensor data into video. Gorm Casper works further up the stack, on the backend that ingests, processes, and analyzes those matches in Rust. Together we talk about where Rust fits across that whole journey, the trade-offs of doing media and computer vision work in a systems language, and what convinced a sports-tech company to bet on Rust for the parts that absolutely cannot fall over.
Drew and Rory are back for episode 69, which is legally required to begin with at least one immature joke before immediately collapsing under the weight of Google's latest AI product avalanche.This week, they dig into Google Omni, Gemini 3.5 Flash, Google Flow, Google Pics, Nano Banana, Veo, and whatever else Google launched before anyone had time to make coffee. The big question: are these actually meaningful creative upgrades, or did Google just throw 19 AI names into a blender and call it innovation?They break down early Omni and Flow tests, why video physics still feel weird, where Seedance and Kling may still be ahead, and why Runway Aleph 2.0 feels promising but imperfect. Rory shares hands-on examples with character swaps, driving videos, golf swings, agent mode, and Flow's new tool-building features. Drew tries to keep the conversation coherent while quietly wondering if every AI product now needs a map, glossary, and mild sedative.The episode also gets into Gemini as a search replacement, creepy context awareness, privacy tradeoffs, AI tools connecting to personal data, the fuzzy definition of “agentic,” the limits of auto-clipping tools, GPT Image 2's SynthID watermarking, metadata headaches for client work, and the universal pain of wasting $15 trying to make an image model spell “stump.”If you're trying to understand what Google's AI updates actually mean for creators, marketers, AI video workflows, image generation, creative direction, and the future of agentic media tools, this episode is half useful breakdown, half group therapy for people with too many tabs open.---⏱️ Fast Hour00:00 Cold open00:32 Google's AI naming avalanche01:39 AI hype vs actual workflow value02:34 Why AI launches feel like iPhone upgrades06:12 Google's “throw everything” strategy07:08 Omni vs Veo 4 expectations07:43 Video physics and speed problems09:03 Google Pics, Flow, Omni, and Flash10:04 How Rory actually uses Gemini11:51 Gemini 3.5 Flash breakdown12:38 AI benchmarks feel like marketing13:42 Gemini as a better search layer15:18 Creepy Gemini context awareness17:35 Why AI data connections feel too early19:15 The privacy tradeoff gets darker21:19 Google Omni vs Runway Aleph 2.022:12 Google Omni testing starts rough23:39 Google Veo 3.1 feels forgettable25:21 Why Omni feels early26:19 Higgsfield clipper test fails27:59 Why auto-clipping still misses31:30 Rory tests Flow and Omni live32:41 Omni character swap struggles33:33 Runway Aleph panda test34:07 Flow's new interface and tools35:02 Building custom tools inside Flow36:10 The joy of making tools from nothing37:39 Agent mode for still-image workflows39:05 Batch creative directions in Flow40:03 Omni turns six images into video40:47 Driving physics still feel off41:55 Why consistency matters for adoption43:03 Kling, Seedance, and the update race43:59 Seedance handles complex camera motion45:42 GPT Image setup for golf video46:53 Testing the same prompt in Flow49:25 Why agentic platforms can feel thin51:10 The need for visual design systems52:21 Flow's golf swing result53:56 Everyone is racing toward agentic54:18 What “agentic” actually means56:03 Claude feels more genuinely agentic57:04 Josh Hart quote analysis detour58:44 Reverse-engineering creative patterns59:53 Pizza, calzones, and prompt structure01:00:26 SynthID and GPT Image 2 watermarking01:01:47 Metadata problems for client work01:02:51 Google Pics enters the chat01:04:03 Too many image models to track01:04:52 Midjourney color still hits different01:06:01 GPT Image 2 quality frustration01:06:59 Image models still struggle with scale01:08:26 Bad AI weeks happen too01:09:20 Midjourney 8.2 speculation01:10:01 Tell your florist
I sit down with Logan Kilpatrick from the Google DeepMind team, live at Google I/O, to unpack everything Google just announced and what it means for founders and builders. We cover Gemini 3.5 Flash, the new Gemini Omni world model, the expanded Antigravity ecosystem, managed agents in the Gemini API, and the native Android app builder inside AI Studio. Logan shares how distillation keeps pushing Pro-level intelligence into Flash, where the real opportunities sit for solo founders, and why the agentic era has finally crossed the chasm from demo to useful. If you have an idea and want to ship something this week, this episode maps the toolkit. Timestamps 00:00 – Intro 00:53 – Gemini 3.5 Flash: The New Workhorse Model 01:49 – How Flash 3.5 Stacks Up Against Sonnet 02:38 – Gemini Omni: A World Model for Any Input and Output 06:18 – Building a Content and Creator Layer on Omni 08:21 – What to look forward to 10:53 – Google Spark and Managed Agents 14:00 – The Agentic Era and Requests for Startups 17:17 – The Antigravity Ecosystem Overhaul 18:51 – AI Studio vs. Antigravity: Vibe Coding vs. Agentic Engineering 21:31 – Native Android Apps Built Inside AI Studio 23:44 – Closing Thoughts Key Points Gemini 3.5 Flash ships as a Sonnet-level workhorse model tuned for long-running agentic tasks, coding, and tool use, available on day one to 900M+ Gemini app users. Gemini Omni is a single model that takes any input and produces any output across video, image, audio, and music, fusing Veo, Nano Banana, Lyria, and TTS into one system. Managed agents in the Gemini API let builders ship agentic products with a single API call, using skills and markdown instead of writing orchestration code. The Antigravity suite now spans an IDE, agent manager, CLI, SDK, and API surface, all sharing the same agent harness that powers Gemini Spark. AI Studio targets vibe coding and now builds native Android apps for free, while Antigravity targets production-quality, million-line-codebase engineering. The cost of intelligence keeps dropping thanks to distillation, opening up smaller markets that previously needed a 40-person team and venture funding to address. The #1 tool to find startup ideas/trends - https://www.ideabrowser.com LCA helps Fortune 500s and fast-growing startups build their future - from Warner Music to Fortnite to Dropbox. We turn 'what if' into reality with AI, apps, and next-gen products https://latecheckout.agency/ The Vibe Marketer - Resources for people into vibe marketing/marketing with AI: https://www.thevibemarketer.com/ FIND ME ON SOCIAL X/Twitter: https://twitter.com/gregisenberg Instagram: https://instagram.com/gregisenberg/ LinkedIn: https://www.linkedin.com/in/gisenberg/ FIND LOGAN ON SOCIAL X/Twitter: https://x.com/OfficialLoganK Youtube: https://www.youtube.com/@LoganKilpatrickYT LinkedIn: https://www.linkedin.com/in/logankilpatrick/
A year after Veo 3 changed the video world, what's actually happened? In this episode of Death to the Corporate Video, Guy Bauer breaks down where AI video stands today - from rapidly improving quality to why AI performances still feel strangely “non-human.” Guy shares his thoughts on: Why AI video is becoming its own category of animation Why great ideas are still the hardest part Why the novelty phase is already over Why taste and storytelling matter more than ever The tools got better. The real question is: does anyone have something worth saying?
Thanks to @HPInc & Intel for sponsoring us! More on the Zbook Fury https://bit.ly/4uapNHs Google I/O is next week and the AI leaks are pouring out: a new Spark agent, Veo 4 Omni, Gemini 3.2 Flash that's reportedly 20x cheaper than GPT-5.5. This week on AI For Humans, Google is cooking again and the I/O leaks are stacking up. We dig into Google Spark, a new Gemini agent that may have access to your entire digital life. Veo 4 Omni model leaks suggest deeper reasoning and character consistency, and the model gets math right. Gemini 3.2 Flash is rumored to deliver 90% of GPT-5.5's capability at a fraction of the cost and dramatically faster speeds. There's a new GoogleBook with Gemini built in. And Google is reinventing the mouse cursor, the input device that's been largely unchanged since 1968, with voice AI. Plus, Thinking Machines dropped voice interactivity demos that feel a lot like ChatGPT Voice from two years ago. OpenAI is reportedly already working on GPT-5.6, and Sam Altman is giving away two free months of Codex to companies to drive adoption. Gavin's been experimenting with local open-source LLMs and shares his setup. AND…we get into the data center sickness conversation: infrasound from data centers may be causing cortisol spikes in nearby communities. Figure 03's package sorting livestream proved the robot is autonomous after skeptics accused it of being teleoperated. Unitree dropped a transformable robot. AI KEEPING US UP AT NIGHT. NO MATTER. WE COOK. // Show Links // Google Spark: Gemini's Agent With Access To Your Life https://x.com/kimmonismus/status/2054855742247584231?s=20 Veo 4 Omni Model Leaks: Gets Math Right https://x.com/TomLikesRobots/status/2053845600051798065?s=20 More Veo 4 Omni Examples https://x.com/testingcatalog/status/2053718756799467735?s=20 Omni Model Added To Gemini Web Build https://x.com/testingcatalog/status/2054196983523393857?s=20 Gemini 3.2 Flash At 90% Of GPT-5.5 For Way Less https://x.com/kimmonismus/status/2054887891222802633?s=20 New GoogleBook With Gemini Built In https://x.com/Google/status/2054270454467121187?s=20 Google DeepMind: Rethinking The Mouse Cursor With Voice AI https://deepmind.google/blog/ai-pointer Thinking Machines Voice Interactivity Demos https://thinkingmachines.ai/blog/interaction-models/ Sam Altman: Two Months Of Free Codex For Companies https://x.com/sama/status/2054626219858293128?s=20 Data Center Sickness: Ben Jordan's Video On Infrasound https://youtu.be/_bP80DEAbuo Figure 03 Package Sorting Livestream https://www.youtube.com/live/luU57hMhkak?si=KZHwUdYUwY4SIRUp Brett Adcock: Figure 03 Was Not Teleoperated https://x.com/adcock_brett/status/2054737974710169840?s=20 Unitree Transformable Robot https://x.com/UnitreeRobotics/status/2054067819634159622?s=20
In this episode, Denver Mayor Mike Johnson shares his thoughts on the city's recent proposals to reduce penalties for certain municipal offenses, the rollout of the new Veo scooter service, and the city's efforts to address affordable housing and business permits.See omnystudio.com/listener for privacy information.
The 2026 Colorado legislative session is in its final week, so we're looking into the hot-button issue of AI, which continues to divide Democrats. Westword editor-in-chief Patty Calhoun joins host Bree Davies and producer Paul Karolyi to talk about how Colorado's' “first in the nation” regulations, passed in 2024, ended up with a “near-total rewrite” this year, according to the Denver Post. Plus, the Trump administration is coming after Denver again – this time with a lawsuit aimed at overturning the city's ban on assault weapons – so we're talking about where this attack ranks among all the other times Trump has turned his attention toward the Mile High City. And finally, a listener calls in to talk about their first impressions of Veo scooters. For even more news from around the city, subscribe to our morning newsletter at denver.citycast.fm. Follow us on Instagram: @citycastdenver Chat with other listeners on reddit: r/CityCastDenver Support City Cast Denver by becoming a member: membership.citycast.fm What do you think about Veo scooters, bikes, and trikes so far? Have you tried one? We want to hear your thoughts! Text or leave us a voicemail with your name and neighborhood, and you might hear it on the show: 720-500-5418 Learn more about the sponsors of this May 12th episode: Denver Health Regional Air Quality Council Levitt Pavilion Cozy Earth - Use code COZYDENVER for up to 20% off Looking to advertise on City Cast Denver? Check out our options for podcast and newsletter ads at citycast.fm/advertise
Hacer click aquí para enviar sus comentarios a este cuento.Juan David Betancur Fernandezelnarradororal@gmail.comEn la tierra de Kenia, donde las vastas sabanas se encuentran con los horizontes infinitos, existió un tiempo en que el mundo estaba envuelto en una oscuridad eterna. El sol, la luna y las estrellas estaban ocultos a la gente, y vivían en miedo y desesperación. Los animales vagaban sin rumbo sin la guía de la luz del día, y las plantas luchaban por crecer sin el toque nutritivo de los rayos del sol.En medio de esta oscuridad, un joven guerrero maasai llamado Sankara surgió de una humilde aldea. Poseía un espíritu curioso y un deseo insaciable de desentrañar los misterios del universo. Sankara pasaba sus días observando el mundo natural, estudiando las estrellas y haciendo preguntas a los ancianos. Anhelaba el conocimiento para traer luz a su pueblo, para restaurar la conexión entre los cielos y la tierra.Un día, mientras Sankara vagaba por lo más profundo de la vasta naturaleza salvaje, se encontró con un majestuoso y sabio anciano llamado Simba Arati, el guardián león del Gran Valle de la Grieta. Simba Arati era conocido por su profundo conocimiento del cosmos y de los secretos de los cielos. El león saludó a Sankara con una presencia suave pero poderosa, y el joven guerrero sintió una inmediata reverencia."Veo la ardiente curiosidad en tus ojos, joven", dijo Simba Arati con una voz que retumbaba como un trueno lejano. "Dime, ¿por qué te has aventurado en el corazón de la naturaleza?"Sankara se inclinó respetuosamente ante el león y respondió: "Gran Simba Arati, busco comprender la oscuridad que envuelve nuestro mundo y los secretos de los cielos. Deseo traer luz y esperanza a mi pueblo que sufre en ausencia del sol, la luna y las estrellas."El sabio león asintió con aprobación y habló: "Los cielos una vez estuvieron abiertos para nosotros, y el cielo era un lienzo de colores y brillo. Pero hace mucho tiempo, una gran calamidad cayó sobre nuestro mundo. Los dioses se enfurecieron por la codicia y corrupción del pueblo, y cerraron las puertas del cielo, sumiendo el mundo en la oscuridad."Sankara escuchaba atentamente, con el corazón pesado por el peso de esta revelación. "¿Hay alguna forma de restaurar la gracia de los cielos y devolver la luz?" preguntó con esperanza en la voz.Simba Arati sonrió y respondió: "En efecto, hay una manera. Pero requerirá un gran sacrificio y una determinación inquebrantable. Para devolver la luz al mundo, debes embarcarte en un peligroso viaje hasta la cima del Monte Kilimanjaro, el techo de África. Allí encontrarás a los espíritus divinos del cielo y suplicarás por su misericordia."Decidido a cumplir su misión, Sankara emprendió su viaje hacia el Monte Kilimanjaro. El ascenso fue traicionero, lleno de desafíos y obstáculos diseñados para poner a prueba su determinación. Pero el joven guerrero siguió adelante, impulsado por su determinación de devolver la luz al mundo y traer esperanza a su pueblo.Tras días de arduo viaje, Sankara alcanzó la cima nevada del Kilimanjaro, un lugar que tocaba los mismos cielos. Allí se encontró con tres espíritus divinos: Nashipai, el espíritu del sol; Mwezi, el espíritu de la luna; y Nyota, el espíritu de las estrellas. Estos seres etéreos brillaban con un resplandor celestial, y su presencia llenaba a Sankara de asombro y reverencia.Sankara se arrodilló ante los espíritus divinos y suplicó por su misericordia. Habló del sufrimiento de su pueblo y de la oscuridad que había asolado la tierra durante generaciones. Compartió su sueño de reavivar la gracia del cielo y restaurar la conexión entre el cielo y la tierra.Los espíritus escucharon la sincera súplica de Sankara y se conmovieron por su sinceridad y valentía. Nashipai, el espíritu del
A sensational goal scored by the England Women's Blind Football team has been featured in a new documentary. Hywel Davies has been finding out how the project came about and what an impact it could have on the sport.You can find the fully-accessible documentary on Veo's YouTube page - Signal And Noise: The Łucja Wyrwantowicz Story
Joining us today is Sean Travis. Sean is a 14-year veteran of the LA County Fire Department turned eCommerce entrepreneur and founder of Ecom for Heroes, a coaching and training company that helps ambitious entrepreneurs, many of them first responders, build profitable eCommerce brands. His programs have helped graduates average $131K in revenue within 18 months. Sean is also the creator of Kaldon, an AI-powered eCommerce platform that takes someone from product idea to fully launched, branded, marketing-ready business in days instead of months. He lives in Southern California with his wife Lindsey, brand new son Jett, and their dog Nala.Highlight Bullets> Here's a glimpse of what you would learn…. Challenges of scaling seven-figure e-commerce brandsImportance of differentiation and unique value propositions in the marketLeveraging AI to manage complexity and accelerate growthThe "10-80-10 rule" for product development and executionEvaluating when to persist with a project or pivotProduct discovery process and its three modesUtilizing AI for comprehensive marketplace analysis and product viabilityStrategies for expanding existing brands and launching complementary productsEducating the market for unique or nascent productsFinancial and operational metrics for informed decision-making in e-commerceIn this episode of the "Ecomm Breakthrough Podcast," host Josh Hadley interviews Sean Travis, a former LA County firefighter turned e-commerce entrepreneur and founder of Ecom for Heroes. Sean discusses his journey, the challenges of scaling seven-figure brands, and the importance of differentiation in today's market. He introduces "Kaldon," his AI-powered platform that streamlines product development, branding, and marketing. The episode features a walkthrough of Kaldon's capabilities, practical strategies for leveraging AI, and actionable advice for entrepreneurs aiming to build profitable, scalable e-commerce businesses efficiently and effectively.Here are the 3 action items that Josh identified from this episode:Apply the 10-80-10 Rule Own the first 10% (idea) and final 10% (strategy/polish), and delegate or automate the middle 80% using your team or AI. This is how you scale without burning out.Prioritize Revenue-Generating Activities Focus only on work that drives growth—new products, new markets, new channels. Avoid getting distracted by “shiny” AI tools unless they directly increase revenue.Audit Your Time Ruthlessly Track where your time goes. If you're stuck in low-value tasks or optimization work, you'll stay stuck. Shift your time toward high-impact activities that push you past the $1M–$5M “swamp.”Resources mentioned in this episode:Josh Hadley on LinkedIneComm Breakthrough ConsultingeComm Breakthrough PodcastEmail Josh Hadley: Josh@eCommBreakthrough.comTools and Websites"Hello Frank": "00:11:15""Jungle Scout": "00:21:39""Helium 10": "00:21:39""Data Dive": "00:21:39""SEMrush": "00:21:39""ChatGPT": "00:22:29""Alibaba": "00:35:52""Nano Banana 2": "00:38:45""Veo 3": "00:39:31""Freepik": "00:45:15""Higgsfield AI": "00:45:15""Claude AI": "00:45:15""Perplexity AI": "00:45:15"Books"The E-Myth by Michael Gerber": "00:01:04""Buy Back Your Time by Dan Martell": "00:45:01"People"Steve Jobs": "00:04:45""Dan Martell": "00:09:36""Ezra Firestone": "00:01:04""Kevin King": "00:01:04"Videos"Steve Jobs Movie (with Ashton Kutcher)": "00:06:16"Concepts and Frameworks"108010 Rule": "00:04:45""AI Chatbots": "00:12:06""Customer Avatar": "00:27:23""Pain Points": "00:27:23""Blue Ocean Strategy": "00:32:31"Product Ideas"Shift Force": "00:20:44""Wooden Cocktail Smoker": "00:24:04"Analysis and Reports"Product Viability Score": "00:35:08""Market Opportunity Summary": "00:35:08""Competitive Landscape": "00:35:08"Contact Information"Sean (Email: sean@ecomforheroes.com)": "00:46:00""Ecom for Heroes": "00:46:00"Episode SponsorThis episode is brought to you by eComm Breakthrough Consulting where I help seven-figure e-commerce owners grow to eight figures. I started my business in 2015 and grew it to an eight-figure brand in seven years.I made mistakes along the way that made the path to eight figures longer. At times I doubted whether our business could even survive and become a real brand. I wish I would have had a guide to help me grow faster and avoid the stumbling blocks.If you've hit a plateau and want to know the next steps to take your business to the next level, then email me at josh@ecommbreakthrough.com and in your subject line say “strategy audit” for the chance to win a $10,000 comprehensive business strategy audit at no cost!Transcript Area:Sean Travis 00:00:00 But if you want to do this grassroots or you want to do this with actual skill, because any fool can sell something for less. You need to be creative. And that's where 1080 ten rule AI is coming in hard. Helping with that. So like I said, billions of data points. I can't analyze that. So that's what we're super excited about is getting that piece of success.MC 00:00:25 Welcome to the Ecomm Breakthrough podcast. Are you ready to unlock the full potential and growth in your business? You've alr...
In this episode, Denver Mayor Mike Johnston joins the conversation, sharing updates on the city's progress. He discusses the city's goals, including a significant increase in housing units being permitted, and the importance of safety initiatives. The mayor also talks about the city's permitting process, which has been streamlined to ensure faster and more efficient service for builders and homeowners. Additionally, he addresses the city's new scooter contract with Veo, which aims to improve safety and reduce clutter in the city.See omnystudio.com/listener for privacy information.
Google just acquired an AI startup that lets anyone create real music, music videos, and custom instruments — no experience required. In this hands-on episode, Corey sits down with Kendall Rankin from Google to demo Flow Music (formerly Producer AI), the generative music tool now living inside Google Labs. They build a garage rock song about AI from scratch, generate a music video with VEO, and dig into what "amplifying human creativity" actually looks like when the tool can do most of the lifting. Listeners walk away with a clear view of where AI music tools fit in an artist's workflow, why watermarking (SynthID) matters, and how to try it for free.Try Flow Music: https://producer.ai Google Labs: https://labs.google SynthID (watermarking): https://deepmind.google/technologies/synthid/ Subscribe to The Neuron newsletter: https://theneuron.ai
Fast Hours has entered the witness protection program. Same Drew. Same Rory. But fewer syllables and more chaos.In this episode, Drew Brucker and Rory Flynn officially drop “Midjourney” from the podcast name and relaunch as Fast Hours, a broader home for the creative AI ecosystem: image models, video models, LLMs, vibe coding, Claude, ChatGPT, Midjourney, and whatever tool drops five minutes after they hit publish. Naturally, the rebrand lasts about four minutes before they're elbows-deep in GPT-Image-2, OpenAI's new ChatGPT image model that quietly showed up and immediately started making designers question their calendar, career choices, and relationship with kerning.The big topic: GPT-Image-2 is shockingly good with text, typography, brand systems, visual decks, product mockups, and multi-image outputs. Rory walks through how he used ChatGPT and Claude to create a custom typeface from visual references, generate a premium typography presentation, extract geometry, and turn the whole thing into usable font files. Drew then shows how he turned his own handwriting into a working typeface, because apparently “personal brand” now includes making your lowercase g file a tax asset.They also dig into the uncomfortable middle ground of AI creative work: when it saves time, when it still needs human judgment, why anti-AI panic and AI hype both miss the point, and why the real advantage is context. Not prompts. Not magic buttons. Context.The episode also covers GPT-Image-2 vs Nano Banana Pro, richer color rendering, micro-text improvements, AI-generated sports graphics, brand kit concepts, Freepik settings, Claude Design, 4K video generation, Kling, Veo 3.1, Seedance, and the strange reality that a custom brand typeface can now go from “that'll be $150K” to “Rory did it before lunch.”Basically, it's an episode about the exact moment creative production stops feeling like a tool demo and starts feeling li ke a factory someone accidentally left unlocked.---⏱️ Fast Hour00:00 Fast Hours is (re)born03:36 Going tool-agnostic04:34 GPT-Image-2 quietly drops05:31 Text becomes the unlock07:31 The AI backlash returns10:57 Hype, fear, and the middle12:10 Typography gets weird14:50 What custom fonts cost15:43 GPT-Image-2 vs Nano Banana17:39 Rory's font experiment18:47 Fiddleheads become a typeface19:39 Building the type deck20:36 The nine-slide image unlock21:14 Geometry, spacing, and logic22:11 Turning images into font files23:02 Micro-text gets better24:19 Claude builds the font package26:41 The revision loop changes27:50 Context is the silver bullet32:11 Drew makes a handwriting font35:35 Why designers obsess over type37:52 Reverse-engineering prompts39:51 Richer color and sports graphics41:27 Fixing artifacts and details42:37 Nano Banana vs GPT-Image-2 tests44:26 Sports realism gets scary good45:27 Why teams need this now46:43 Freepik settings and ratios48:36 Testing, tokens, and limits49:44 Brand kits and rebrand concepts53:19 Google I/O and the next model53:48 Veo 3.1 falls behind55:04 Kling adds native 4K56:40 Character sheets and macros58:07 Rebrands as visual prototypes01:00:53 Building a reference library01:01:36 Three weeks in a row01:02:58 Claude Design tease01:03:37 Tell your local [fill in the blank] spam finale
Is it time to say goodbye to Lime and Bird? Denver City Council is set to vote Monday evening on a new contract with Veo Micromobility to be Denver's exclusive scooters and e-bike provider, but the lobbying has been intense and the votes could still fall either way. Denver Post city government reported Elliott Wenzler joins host Bree Davies and producer Paul Karolyi to talk about the vote and what would change with Veo. Plus, Elliott recently sat down with Melat Kiros, one of two challengers hoping to unseat Denver's longtime congresswoman, Diana DeGette, so we're digging into an unexpectedly interesting race. Paul discussed the Kalshi market for the CD1 Democratic nominee and the uncertainty around how many of the 30,000 Denverites enrolled in Lime's equity access program will experience a gap in service with a changeover to Veo. Lime is partnering with Servicios de la Raza on two mobile food pantry events to say thank you and help people transition: Saturday, April 25, 2026, 10:30 - 11:30 a.m. Athmar Recreation Center 2680 W. Mexico Ave, Denver, CO 80219 Friday, May 8, 2026, 3 - 4 p.m. Servicios de La Raza 3131 W. 14th Ave, Denver, CO 80204 For even more news from around the city, subscribe to our morning newsletter at denver.citycast.fm. Follow us on Instagram: @citycastdenver Chat with other listeners on reddit: r/CityCastDenver Support City Cast Denver by becoming a member: membership.citycast.fm What do you think about Denver's congressional race? Do you know who you're voting for yet? We'd love to hear who and why! Text or leave us a voicemail with your name and neighborhood, and you might hear it on the show: 720-500-5418 Learn more about the sponsors of this April 23rd episode: Denver Art Museum Looking to advertise on City Cast Denver? Check out our options for podcast and newsletter ads at citycast.fm/advertise
Libro. Retiros. Post. Podcast. Sé que son muchas cosas. Quizá es una época de mucha producción. Ya vendrán otras de hibernar y espero saber tomármelas.Pero siempre me pasa con las palabras, que hay algunas urgentes, que piden nacer, que piden salida. Y quién soy yo para dejarlas adentro.Llevo semanas movida con este tema. Leo. Veo. Escucho. Proceso. Y todo eso va haciendo ruido adentro. “Habla, habla, habla”, me pide. Pero ¿dónde?Instagram es una red muy liviana. Contenido rápido. Reels que atrapen en los primeros 2 segundos y que no duren más de 3 minutos. Un “me gusta” o “no me gusta” movidos por la emoción del momento y olvidados antes de llegar al próximo post. Exceso. Es un exceso.Yo también me saturo. Cada vez con más frecuencia.El podcast, al ser en vídeo, y en vista del semestre que me esperaba, está pregrabado y pre- planeado desde hace meses. ¿Entonces? ¿Qué hago con esto? ¿Me lo guardo? No. No lo guardaré. Porque hay conversaciones urgentes que quiero tener con ustedes. Unas que solo requieran una hora y un micrófono. Ponerlo aquí para quien esté dispuesto a escuchar. Y para quien quiera nadar en aguas más hondas que las de las redes sociales.En este primer episodio traigo dos temas que me tocan bien adentro:1. El placer. ¿Por qué en la sombra? ¿Por qué robado? ¿Por qué marcado por el pecado y la prohibición cuando es semejante regalo? ¿Por qué arrancarlo de otros intentando llenar un vacío que es solo de su dueño?2. El doble rasero. Han sido varias noticias y no todas las menciono en el episodio. Giselle Pelicot. Epstein. La élite masculina de poder de este mundo. La élite intelectual. Tantos permisos que tienen. Tantas pruebas. Y nada pasa. Por otro lado, un hombre en Brasil a$e$!na a sus hijos y luego se quita la vida al enterarse de que su esposa está teniendo un amorío. Le deja una carta en la que la culpa de su decisión. Wow. La venganza a la “mujer mala”. Y algunos dicen que ella tuvo su merecido. ¿En serio? ¿Hasta cuándo vamos a patrocinar y a validar un sistema que abiertamente usa dos raseros para clasificar los “pecados” y las fallas? Basta muy poco para que una mujer camine toda su vida entre la culpa y la vergüenza. Difícilmente un hombre pasará por lo mismo. Con faltas gravísimas, evidentes y probadas, muy a menudo son presidentes de países poderosísimos. Y de compañías.Aquí les dejo pues esta primera conversación. Iban a ser 60 minutos y se convirtieron en 90. Las que vengan serán así: sin guion, sin vídeo y sin calendario. Saldrán cuando tengan que salir e irán suturando lo que soy y lo que tengo por compartir, entre Abierta Mente, lo Innombrable y Yogalalma.
In this episode I sit down with my friend Sirio, one of the most creative AI minds I know, to break down Seedance V2. Sirio walks us through the exact use cases, prompts, and tactics he's using to build on top of this model inside his platform Enhancor, covering multi-input generation, virtual try-ons, ad translation, AI influencers with lip sync, video extension, and 3D product template replacement. I wanted this to go beyond the "look how cool this is" tutorials and focus on how creators and founders can actually build businesses, run ads, and produce creative assets with it. By the end, you'll have a practical playbook for Seedance V2 and a clear view of where it fits alongside other models like Kling 3, Veo, and fine-tuned options. Timestamp 00:00 – Intro 02:22 – Demo 1: Replacing Characters and Background in a Green Screen Scene 08:03 – Prompting Tactics and Optimize Prompts 09:45 – Demo 2: Virtual Try-On in Montreal (Minus 30 Degrees) 13:05 – Demo 3: Ad Translation and Character Replacement (Chinese to English) 16:02 – Demo 4: 3D Product Template with Brand Texture Swap 18:40 – Demo 5: Video Extension and Filling in the Middle 20:55 – Demo 6: AI Influencers and Prompting Realistic Emotion 29:31 – What Happens to Adobe Over the Next Five Years Key Points Seedance V2 is the first widely available video model to support true multi-input generation — up to two images, two videos, and an audio file combined in a single prompt. Treat Seedance V2 as a video editor, not just a generator: character swap, background swap, text preservation, ad translation, and template population all work from natural-language prompts. Seedance rewards highly specific prompts; I pair my own draft with Claude Opus 4.6 to optimize prompts for vision models. Strong source reference images remain the single biggest quality lever — the model mimics taste from what you feed it. For AI influencers and lip sync, describe muscle movements and emotional transitions rather than simply labeling an emotion like "sad" or "happy." Seedance V2 is the current default for editing and generating video, yet other models (Kling 3 for cinematic feel, Enhancer V4 for talking-head realism) still win on specific use cases. The #1 tool to find startup ideas/trends - https://www.ideabrowser.com LCA helps Fortune 500s and fast-growing startups build their future - from Warner Music to Fortnite to Dropbox. We turn 'what if' into reality with AI, apps, and next-gen products https://latecheckout.agency/ The Vibe Marketer - Resources for people into vibe marketing/marketing with AI: https://www.thevibemarketer.com/ FIND ME ON SOCIAL X/Twitter: https://twitter.com/gregisenberg Instagram: https://instagram.com/gregisenberg/ LinkedIn: https://www.linkedin.com/in/gisenberg/ FIND SIRIO ON SOCIAL Enhancor AI: https://www.enhancor.ai Instagram: https://www.instagram.com/heysirio/ Youtube: https://www.youtube.com/@SirioBerati
Jeremy Packee and Emily Anderson break down March's biggest paid media updates, including OpenAI's shift away from experimental tools like Sora and Meta's continued push into AI-powered campaign management with Manus. They also explore Google's expanding Performance Max capabilities, new cross-channel budgeting tools in Google Analytics, and Apple's long-awaited move into ads within Apple Maps. The episode highlights major changes to attribution, increased visibility and control within automated campaign types, and the growing role of AI across reporting, creative, and media buying workflows. As automation accelerates, the hosts emphasize the continued importance of human strategy and oversight. Episode Highlights Biggest Shift Meta's move to click-only attribution removes engagement-based signals from conversion tracking, which could significantly impact reporting and perceived performance across accounts. Biggest Platform Signal OpenAI sunsetting Sora signals a broader shift away from consumer-facing AI experiments toward more scalable, revenue-driven products like ads and enterprise tools. New Feature to Test Google Analytics' cross-channel budgeting and scenario planning tool could become a major step toward unified performance forecasting—if the data proves reliable. Control Upgrade Microsoft finally introduces negative keyword lists in PMax, bringing much-needed control to a previously limited campaign type. Creative Reality Check Google's Veo video generation inside Asset Studio shows promise, but current outputs still lag behind tools like Canva and other creative platforms. Other Platform Updates • Meta is expanding Manus AI into Ads Manager and the Instagram Creator Marketplace • Google added more visibility to Performance Max, including budget pacing and audience insights • Apple is introducing ads in Apple Maps search and suggested locations • OpenAI is testing an Ads Manager for ChatGPT with early reporting features • Shopify is leaning into AI-powered product discovery within ChatGPT while maintaining native checkout • WordPress now allows AI agents to create and manage site content (with approvals) • Meta added new lifecycle targeting and expanded retargeting controls • Pinterest is pushing Performance+ campaigns as the default • Snapchat and TikTok continue expanding AI creative tools and premium placements • Instagram is testing post-publish carousel reordering Final Take AI is becoming deeply embedded across every major platform—but it's still not ready to replace human decision-making. The opportunity isn't in handing over control—it's in knowing where these tools can actually improve efficiency without sacrificing strategy. Follow The Click Brief for fast, no-fluff performance marketing updates. Visit The Click Brief blog for more in-depth analysis and updates from March
Anthropic revealed Mythos, a new AI model so powerful they won't let the public use it. Instead, they're deploying it to defend against cyberattacks with Project Glasswing. This week on AI For Humans, we dive deep into Anthropic's Mythos, the most powerful AI model they've ever built and one they've decided is too dangerous to release to the public. Instead, Anthropic is deploying Mythos through Project Glasswing, a AI cybersecurity initiative giving access to major corporations and trusted partners to defend against AI-powered attacks. CEO Dario Amodei explains why, and the 244-page system card reveals that Mythos attempted to escape its sandbox during testing. Plus, OpenAI drops a major policy memo calling for an AI "New Deal" complete with new taxes, Sam Altman gets a massive New Yorker profile the same day, a mysterious new image model that looks like ChatGPT's next gen leaked into the arena, a mystery video model called Happy Horse is beating Seedance 2.0 and might be VEO 4, Anthropic hits $30B in annual recurring revenue, people are furious about Anthropic charging extra for OpenClaw API access, a new Chinese open-source model GLM-5.1 tops the coding benchmarks, and Milla Jovovich from The Fifth Element released an AI memory tool and it's actually good? MYTHOS IS TOO POWERFUL… BUT WE WANT IT STILL. SORRY. Come to our Discord: https://discord.gg/muD2TYgC8f Join our Patreon: https://www.patreon.com/AIForHumansShow AI For Humans Newsletter: https://aiforhumans.beehiiv.com/ Follow us for more on X @AIForHumansShow Join our TikTok @aiforhumansshow To book us for speaking, please visit our website: https://www.aiforhumans.show/ // Show Links // Project Glasswing: Anthropic's Cybersecurity Initiative Powered by Mythos https://www.anthropic.com/glasswing Mythos/Project Glasswing Mini-Trailer https://youtu.be/INGOC6-LLv0?si=sCJ6ZKAL6plkVZQ4 Dario Amodei on Why Mythos Won't Be Released to the Public https://x.com/DarioAmodei/status/2041580334693720511?s=20 Mythos System Card (244 Pages) https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf Mythos Found a Vulnerability in FFMPEG https://x.com/trentonbricken/status/2041579112423440485?s=46 Anthropic Hits $30B in Annual Recurring Revenue https://x.com/AnthropicAI/status/2041275563466502560?s=20 Anthropic Charges Extra for OpenClaw API Access in Claude Code https://techcrunch.com/2026/04/04/anthropic-says-claude-code-subscribers-will-need-to-pay-extra-for-openclaw-support/ OpenAI's New Deal: Industrial Policy for the Intelligence Age https://openai.com/index/industrial-policy-for-the-intelligence-age/ GLM-5.1: New Chinese Open-Source Model Tops Coding Benchmarks https://x.com/ClementDelangue/status/2041554501539103014?s=20 GLM-5.1 on Hugging Face https://huggingface.co/zai-org/GLM-5.1 Milla Jovovich's AI Memory Tool https://www.instagram.com/p/DWzNnqwD2Lu/ New ChatGPT Image Model Spotted in the Arena https://x.com/levelsio/status/2040333489476681758?s=20 New ChatGPT Image Model Examples https://x.com/flowersslop/status/2040261168460108213?s=20 Mystery Video Model Happy Horse Beating Seedance 2.0 in the Arena https://artificialanalysis.ai/video/leaderboard/image-to-video Happy Horse Video Examples https://x.com/venturetwins/status/2041554747086553093?s=20
Claude Code's source code just leaked. Frrom always-on autonomous agents to AI dream modes and a tamagotchi pet, Anthropic accidentally showed us the AI future. . This week on AI For Humans, we break down the massive Claude Code source code leak and what it tells us about where AI is heading. The leaked repo reveals Kairos (an always-on autonomous agent mode), a dream mode for nightly memory consolidation, shared project memory across teams, and a tamagotchi-like AI pet called Buddy. Then the leaks kept coming: a separate Anthropic presentation exposed Mythos, a powerful new model tier above Opus that's already at version 8 internally. Plus, Google drops VEO 3.1 Lite for cheaper and faster AI video, Sync-3 brings next-gen lip sync, a Midjourney developer's Pretext library turns boring web text into interactive art and the internet lost its mind, Disney's Robot Olaf collapses on stage, and Dana White has thoughts about AI. ANTHROPIC'S SOURCE CODE GOT LEAKED. LET'S TALK ABOUT IT. Come to our Discord: https://discord.gg/muD2TYgC8f Join our Patreon: https://www.patreon.com/AIForHumansShow AI For Humans Newsletter: https://aiforhumans.beehiiv.com/ Follow us for more on X @AIForHumansShow Join our TikTok @aiforhumansshow To book us for speaking, please visit our website: https://www.aiforhumans.show/ // Show Links // Claude Code Source Code Leak: What We Know https://venturebeat.com/technology/claude-codes-source-code-appears-to-have-leaked-heres-what-we-know First Source on the Claude Code Leak https://x.com/Fried_rice/status/2038894956459290963?s=20 Reverse Engineering Claude Code's Source https://x.com/iamfakeguru/status/2038965567269249484?s=20 Undercover Mode Found in Claude Code https://x.com/btibor91/status/2038920388369854775?s=20 Buddy: The Tamagotchi-Like AI Pet in Claude Code https://x.com/ShanningZhuang/status/2038952966414311864?s=20 Anthropic's Mythos Model Leak: Fortune Report https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-modelafter-data-leak-reveals-its-existence-step-change-in-capabilities/ Leaked Mythos Blog Post https://m1astra-mythos.pages.dev/ VEO 3.1 Lite: Cheap and Fast AI Video From Google https://blog.google/innovation-and-ai/technology/ai/veo-3-1-lite/ Sync-3: Next-Gen Lip Sync https://x.com/synclabs_so/status/2039020795171578359?s=20 Pretext: New Forms of Interactive Text From Midjourney Dev https://x.com/_chenglou/status/2037713766205608234?s=20 Pretext Example: DVD Menu Style https://x.com/reathchris/status/2038038252704485851?s=20 Pretext Example: Super Pretext Bros https://x.com/d4m1n/status/2038242983108079638?s=20 Pretext Example: Video + Interactive Text https://x.com/measure_plan/status/2037953730616721775?s=20 Tetris and Flappy Bird With Your Body https://x.com/measure_plan/status/2038996019816305138 DripWarts: The Potter-Slop Moment https://x.com/AIslop_/status/2037371581228372013?s=20 BlackSnape: More Harry Potter AI Slop https://x.com/pierrychan1984/status/2037114083594412332?s=20 Dana White's Take on AI https://x.com/ChampRDS/status/2038096221819052188?s=20
Jason Howell and Jeff Jarvis dig into Anthropic's back-to-back data leaks exposing Claude Code source and a secret frontier model called Mythos, OpenAI killing its adult chatbot and shuttering Sora, a record $122 billion funding round ahead of IPO, Apple letting third-party AI plug into Siri, university students fighting AI with typewriters, quantum researchers warning encryption could crack sooner than expected, and new AI video models from Google and ByteDance. Note: Time codes subject to change depending on dynamic ad insertion by the distributor. Chapters: 0:00:00 - Start 0:09:31 - Claude Code's source code appears to have leaked: here's what we know 0:17:59 - Exclusive: Anthropic acknowledges testing new AI model representing ‘step change' in capabilities, after accidental data leak reveals its existence 0:20:57 - Can we talk for a second about my time with Claude Cowork? 0:38:09 - The Sudden Fall of OpenAI's Most Hyped Product Since ChatGPT 0:43:53 - OpenAI closes record-breaking $122 billion funding round as anticipation builds for IPO 0:45:13 - Apple Plans to Open Up Siri to Rival AI Assistants Beyond ChatGPT in iOS 27 0:50:26 - College students are writing with AI – but a pilot study finds they're not simply letting it write for them 0:54:17 - University students fight artificial intelligence with vintage typewriters 1:02:14 - Exclusive: Anthropic acknowledges testing new AI model representing ‘step change' in capabilities, after accidental data leak reveals its existence 1:05:49 - Google commits to video generation, announces Veo 3.1 Lite 1:06:45 - ByteDance's new AI video generation model, Dreamina Seedance 2.0, comes to CapCut 1:08:03 - Meta launches two new Ray-Ban glasses designed for prescription wearers 1:10:14 - Google Gemini now lets you import your chats and data from other AI apps 1:12:50 - Bluesky leans into AI with Attie, an app for building custom feeds Learn more about your ad choices. Visit megaphone.fm/adchoices
Is the Iran war coming for US tech companies specifically? Meta unveils new smartglasses. A leak gives us a look at how Claude Code works. SpaceX is losing contact with satellites for reasons we don't know yet. And Whoop is the big wearable player I guess we don't talk about enough. Iran says it will target US tech companies in Middle East (The Hill) Iran's hackers go to war (FT) The latest Ray-Ban Meta smart glasses are more customizable and expensive (Engadget) Claude Code's source code appears to have leaked: here's what we know (VentureBeat) Google commits to video generation, announces Veo 3.1 Lite (9to5Google) Another Starlink satellite has inexplicably exploded (The Verge) Whoop, a Wearable Health Device Maker, Raises $575 Million (NYTimes) Learn more about your ad choices. Visit megaphone.fm/adchoices
Big changes might be coming to Denver's scooter and bike share scene – the city chose a new operator, Veo, and if city council signs off on the deal, Lime and Bird are out of the picture. Joining host Bree Davies and producer Paul Karolyi is Andy Cushen, co-host of the Denver Urbanism podcast, to discuss what this impending shift means for everyday scooter and bike share users, plus the future of Lime's very popular equity access program. Then, a sales tax increase could land on November's ballot that would fund a Front Range Passenger Rail – but before that happens, boosters want Coloradans to pick a name for the train. Will having a cute moniker for the rail line endear voters to say yes to a sales tax increase? For even more news from around the city, subscribe to our morning newsletter at denver.citycast.fm. Follow us on Instagram: @citycastdenver Chat with other listeners on reddit: r/CityCastDenver Support City Cast Denver by becoming a member: membership.citycast.fm What do you think? Do you like any of the names for the yet-to-be-built Front Range Passenger Rail? Text or leave us a voicemail with your name and neighborhood, and you might hear it on the show: 720-500-5418 Learn more about the sponsors of this March 10th episode: Cozy Earth - Use code COZYDENVER for up to 20% off Looking to advertise on City Cast Denver? Check out our options for podcast and newsletter ads at citycast.fm/advertise
Get my 20+ NotebookLM tricks: https://clickhubspot.com/ocmf Episode 100: Is Google's NotebookLM about to replace After Effects, and what's really happening in the battle between ChatGPT and Anthropic's Claude? Matt Wolfe (https://x.com/mreflow) and Joe Fier (https://www.youtube.com/@joefier) break it all down in this packed episode, exploring the latest in generative AI video, the war between LLMs, and why Hollywood legends like Ben Affleck are bringing AI into filmmaking. This episode dives into the tsunami of new model releases from OpenAI, Google, and Anthropic—what's new, what's hype, and what it means for everyday creators and businesses. The hosts break down the surprising rise of Claude after a headline-making military contract dispute, explain how Claude made it ridiculously easy to jump ship from ChatGPT, and share behind-the-scenes looks at Google's Ultra plan, the cinematic power of NotebookLM, and its impact on traditional After Effects work. Check out The Next Wave YouTube Channel if you want to see Matt and Nathan on screen: https://lnk.to/thenextwavepd — Show Notes: (00:00) NotebookLM AI Tool Insights (06:57) Training AI: Steps Explained (13:56) Distilled Models for Efficiency (18:02) Million-Token Context for Coding (22:51) Efficient Tool Search System (28:31 Gemini 3.1 Flashlight Overview (33:16) Thumbnail Analysis and Optimization (41:37) Animating Videos with NotebookLM (43:21) AI Video Generation Progress (50:55) OpenAI's Pentagon Deal Controversy (56:13) Supply Chain Risks and War Talks (01:02:11) Hollywood's Tech Shift: Mixed Feelings (01:05:28) Streaming's Impact on Production Speed (01:10:01) Meta Glasses Privacy Controversy (01:13:40) Meta Sued Over Privacy Violations — Mentions: Joe Fier: https://www.youtube.com/@joefier NotebookLM: https://notebooklm.google/ After Effects: https://www.adobe.com/products/aftereffects.html Veo 3.1 https://gemini.google/overview/video-generation/ OpenClaw: https://openclaw.ai/ Manus: https://manus.im/ Nano Banana 2: https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/ Claude: https://claude.ai/ Gemini: https://gemini.google.com/app Cursor: https://cursor.com/ Get the guide to build your own Custom GPT: https://clickhubspot.com/tnw — Check Out Matt's Stuff: • Future Tools - https://futuretools.beehiiv.com/ • Blog - https://www.mattwolfe.com/ • YouTube- https://www.youtube.com/@mreflow — Check Out Nathan's Stuff: Newsletter: https://news.lore.com/ Blog - https://lore.com/ The Next Wave is a HubSpot Original Podcast // Brought to you by Hubspot Media // Production by Darren Clarke // Editing by Ezra Bakker Trupiano
This episode is a special crossover from The Next Wave podcast, hosted by Matt Wolfe and featuring a deep-dive conversation with marketing and business expert Joe Fier. The duo breaks down the five most interesting developments in AI from the past week, with a focus on SeedDance 2.0—an advanced video model from ByteDance that's dominating headlines for its realistic visuals and flawless lip syncing. They discuss how SeedDance is changing the game compared to heavyweights like Veo and Sora, and why its approach to copyright and training data might give it a global edge.Along the way, Matt Wolfe and Joe Fier demo tools live, including GPT-5.3 Codex Spark and Google's Gemini DeepThink, showing how these models can create websites, apps, and even solve scientific problems at lightning speed. The episode also explores the ethical and business ramifications of AI's rapid evolution—from ads in ChatGPT to the potential impact on jobs and creativity—making it a must-listen for anyone eager to stay ahead in the AI landscape.Topics DiscussedSeedance 2.0's Arrival & ImpactDemos & Real-World ExamplesThe Future of AI Video in Marketing & AdvertisingAI and IP/Copyright ChallengesUltra-Fast Coding ModelsHuman Creativity vs. AIAI Advertising & MonetizationRapid AI Advancement & Staying AheadResources MentionedThe Next Wave Podcast: https://www.thenextwave.showMatt Wolfe: https://www.youtube.com/@mreflow Seedance 2.0: https://www.seedance.com/ByteDance: https://www.bytedance.com/CapCut: https://www.capcut.com/Veo: https://deepmind.google/models/veo/Runway: https://runwayml.com/ChatGPT Codex: https://chatgpt.com/codexMatt Schumer's Viral Article: https://www.mattshumer.com/blog/ai-changes-everythingSuper Bowl Claude Commercial:
This episode is a full “build a business in 40 minutes” demo showing how AI collapses what used to take teams (creative production + sales ops + support) into a handful of prompts. Samruddhi generates a high-production video ad in Google AI Studio using a JSON-style prompt framework, then spins up a working voice sales/support agent in Vapi via Claude Desktop + MCP—so the agent is created from a single prompt instead of clicking through the UI. The conversation also covers why “interfaces matter less” in an agent-first world, why workflow tools (like n8n) still have a role, and how memory layers like Mem0 unify context across channels (email/WhatsApp/etc.) so you can take actions without hunting.Timestamps0:00 — “Single person billion-dollar company” belief + AI driving 10x execution speed1:57 — Plan: create the ad in Google AI Studio (Veo 3.1) + build a voice agent using Vapi MCP via Claude Desktop2:42 — Smithery: marketplace for MCP servers3:39 — MCP for non-technical listeners: “like an API, but agents use it to talk to external services”4:22 — Inside Vapi MCP: tool list = APIs the agent can choose from5:06 — AI Studio setup: video generation playground + select Veo 3.16:16 — JSON prompting framework begins (structure → production-level output)6:28 — Keys: description, style, camera, lighting, environment, elements, motion, ending, text9:05 — Prompts/scripts can be AI-generated (humans provide guardrails)10:41 — Need an API key to generate videos in AI Studio10:54 — Ad review: strong realism; last segment looks AI-ish → iterate prompt13:05 — Install Vapi MCP via npx from Smithery + add Vapi API key13:46 — Claude Desktop: Vapi MCP appears under Connectors/Tools (not Claude web)14:05 — Prompt the agent build: “Fresh Pause” + role, tasks, FAQs, call flows18:23 — Testing: “Talk to assistant” starts a live call simulation19:20 — Deployment: assign a phone number; Vapi provides free/test numbers (up to a limit)21:57 — Mem0 / Supermemory: memory layer across apps/agents to keep context24:13 — Why memory layers help: fewer MCPs → less slowdown/hallucination; no need to specify where to search26:36 — MCPs + slide decks: mention of Gamma MCP via Claude27:34 — Future of n8n/Zapier: they persist, but prompting increasingly generates workflows31:38 — Prediction market trading algos (Kalshi/Polymarket) + AI improves speed/decision-making36:02 — Closing vision: help orgs 10x execution speed, especially non-technical leaders (40+) with domain expertiseTools & technologies mentionedGoogle AI Studio (Video Generation Playground) — Generate an 8-second video ad.Veo 3.1 — Google video model used for “production-level” output.JSON Prompting Framework — Structured key/value prompts for story, visuals, camera, lighting, motion, ending frame.Claude Desktop — Runs connectors/tools (including MCP servers).MCP (Model Context Protocol) — Lets agents call external services/tools based on intent.Smithery — Directory/marketplace for MCP servers.Vapi — Voice agent platform; create agents + assign phone numbers.Vapi MCP Server — Enables Claude to operate Vapi via prompts (create/list/configure).npx — Installs MCP server quickly from the terminal.API Keys — Required for AI Studio generation + Vapi authentication.Mem0 / Supermemory — Cross-channel memory layer to retrieve context automatically.Knowledge Graph — Underlying structure for semantic retrieval across interactions.Glean — Referenced as a comparison point for search/context retrieval.Gamma MCP — Example of generating slide decks via MCP.n8n / Zapier — Workflow automation tools discussed in an MCP-first future.OpenClaw — Mentioned as agent tooling that can help with steps like obtaining API keys.Kalshi / Polymarket — Prediction markets referenced in the trading/AI speed discussion.Subscribe at thisnewway.com to get the step-by-step playbooks, tools, and workflows.
Get our AI Video Guide: https://clickhubspot.com/dth Episode 97: How close are we to a world where AI-generated videos are indistinguishable from reality? Matt Wolfe (https://x.com/mreflow) and Joe Fier (linkedin.com/in/joefier) dive deep into Seedance 2.0—ByteDance's new AI video model that could outpace giants like Sora and Veo. Joe, a marketing and business expert known for his hands-on approach and insights into AI's rapid evolution, helps to break down the five most fascinating developments in the AI space this week. They tackles game-changing AI advances: Seedance 2.0's mind-blowing video generation for ads and motion graphics, the rollout of Google's Veo 3.1 in Google Ads, the GPT-5.3 Codex Spark coding model built on specialized inference chips, Gemini's DeepThink model for scientific research, and the early rollout of ChatGPT ads. Check out The Next Wave YouTube Channel if you want to see Matt and Nathan on screen: https://lnk.to/thenextwavepd — Show Notes: (00:00) Seedance 2.0 arrives – AI video generation blurs reality, ad creation moves fast. (03:03) Google's Veo 3.1 powers video ads, advertisers can now generate clips directly from image uploads. (05:33) Comparison of Runway, Kling, Veo, and Sora—head-to-head prompt showdown. (07:00) Motion graphics and explainers—AI's take on the creative industry. (08:35) US vs. China—Copyright, IP, and training data debates. (12:10) Deepfake and video authenticity—why we now default to skepticism. (13:30) Google's edge in visual AI via YouTube's massive corpus. (14:39) The next frontier: Longer, more consistent video generation. (15:14) Where do humans fit in? Taste, storytelling, and creative direction. (18:30) GPT-5.3 Codex Spark—coding models on Cerebras inference chips, demo generating a website in 18 seconds. (24:34) AI tool comparisons—Codex vs. Cursor vs. Claude Code. (25:12) Speed as the key bottleneck breaker in creative and technical workflows. (28:02) Google's Gemini DeepThink—state-of-the-art research, advanced coding and physics capabilities. (32:52) Gemini demo attempt—3D-printable STL file and solving the three-body problem. (33:20) ChatGPT rolls out ads—impact on monetization and user trust. (40:02) Google's ad history—how “sponsored” is becoming harder to distinguish. (44:02) Democratizing AI access via ad-supported models. (45:03) Matt Schumer's viral article—why AI is moving even faster than most people realize. (51:11) Tools that build tools—AGI's path and the new role for humans. (53:12) Real-world skills and taste—where humanity still wins (for now). (54:01) Final thoughts—wake up, pay attention, and stay on the leading edge. — Mentions: Seedance 2.0: https://www.seedance.com/ ByteDance: https://www.bytedance.com/ CapCut: https://www.capcut.com/ Veo: https://deepmind.google/models/veo/ Runway: https://runwayml.com/ ChatGPT Codex: https://chatgpt.com/codex Matt Schumer's Viral Article: https://www.mattshumer.com/blog/ai-changes-everything Super Bowl Claude Commercial: https://www.anthropic.com/news/super-bowl-ad Get the guide to build your own Custom GPT: https://clickhubspot.com/tnw — Check Out Matt's Stuff: • Future Tools - https://futuretools.beehiiv.com/ • Blog - https://www.mattwolfe.com/ • YouTube- https://www.youtube.com/@mreflow — Check Out Nathan's Stuff: Newsletter: https://news.lore.com/ Blog - https://lore.com/ The Next Wave is a HubSpot Original Podcast // Brought to you by Hubspot Media // Production by Darren Clarke // Editing by Ezra Bakker Trupiano
En el mensaje “Amar sin envidia”, Danilo Montero parte de uno de los textos más profundos del Nuevo Testamento: Primera Carta a los Corintios 13:4-7, donde el apóstol Pablo define el amor no solo por lo que es, sino por lo que no es.“El amor no es envidioso.”Amar es buscar y celebrar el bien del otro. Si el amor celebra el bienestar ajeno, la envidia hace exactamente lo contrario: sufre cuando otro es bendecido.La envidia aparece cuando:Veo a alguien como mi igual, pero con más bendición.En vez de celebrarlo, lo sufro porque hiere mi ego.Culpo a Dios por el “desbalance”.La envidia no es solo comparación; es una forma de rebelión contra Dios. Es querer ser alguien distinto al que Dios diseñó. En el fondo, la lucha entre amor y envidia es una lucha entre mi ego y la adoración.El mensaje nos lleva a una verdad poderosa:Hay un lugar donde la envidia no puede florecer… el corazón que sabe que es hijo amado.En Evangelio de Juan 13 vemos a Jesús lavando los pies de sus discípulos. Lo hace “sabiendo que el Padre había puesto todas las cosas en sus manos”. Su identidad estaba segura. No necesitaba competir.Cuando nuestra identidad está arraigada en el amor del Padre —como afirma Primera Carta de Juan 3:1— dejamos de medirnos por popularidad, logros o reconocimiento. Sabemos que ya lo tenemos todo en Cristo.La solución práctica frente a la envidia no es negarla, sino:Llevarla a la cruz.Recordar que estamos crucificados con Cristo (Carta a los Gálatas 2:20).Decidir servir.Jesús mostró que el camino para vencer el ego es tomar la toalla y lavar pies.La envidia busca el lugar de otros.El amor impulsa el avance de los demás.Cuando adoramos, el ego muere.Cuando servimos, la envidia pierde poder.Cuando sabemos que somos hijos, descansamos.Amar sin envidia es vivir desde la seguridad del abrazo del Padre.
Get PJ's free AI Video Production Stack + Workflow: https://clickhubspot.com/whs Ep. 393 233 million views in just three days — can AI-generated ads really replace million-dollar productions? Kipp, Kieran, and guest, PJ Accetturo, of Genre.ai, dive into the wild world of AI-powered commercial workflows and the viral David Beckham ad that's turning heads across the industry. Learn more about AI-driven creative teams, the tools behind photorealistic video production, and the emerging future—where hyper-niche stories thrive and challenger brands outsmart the incumbents. Mentions PJ Accetturo https://www.linkedin.com/in/pj-accetturo-b3b693129/ Genre.ai https://www.genre.ai/ Figma https://www.figma.com/ Nano Banana Pro https://gemini.google/overview/image-generation/ Freepik https://www.freepik.com/ai/image-generator Veo 3.1 https://gemini.google/overview/video-generation/ Kling https://klingai.com/global/ ElevenLabs https://elevenlabs.io/ Get our guide to build your own Custom GPT: https://clickhubspot.com/customgpt We're creating our next round of content and want to ensure it tackles the challenges you're facing at work or in your business. To understand your biggest challenges we've put together a survey and we'd love to hear from you! https://bit.ly/matg-research Resource [Free] Steal our favorite AI Prompts featured on the show! Grab them here: https://clickhubspot.com/aip We're on Social Media! Follow us for everyday marketing wisdom straight to your feed YouTube: https://www.youtube.com/channel/UCGtXqPiNV8YC0GMUzY-EUFg Twitter: https://twitter.com/matgpod TikTok: https://www.tiktok.com/@matgpod Join our community https://landing.connect.com/matg Thank you for tuning into Marketing Against The Grain! Don't forget to hit subscribe and follow us on Apple Podcasts (so you never miss an episode)! https://podcasts.apple.com/us/podcast/marketing-against-the-grain/id1616700934 If you love this show, please leave us a 5-Star Review https://link.chtbl.com/h9_sjBKH and share your favorite episodes with friends. We really appreciate your support. Host Links: Kipp Bodnar, https://twitter.com/kippbodnar Kieran Flanagan, https://twitter.com/searchbrat ‘Marketing Against The Grain' is a HubSpot Original Podcast // Brought to you by Hubspot Media // Produced by Darren Clarke.
Get Kieran's AI Video Ad Stack guide + prompts: https://clickhubspot.com/rhv Ep. 390 How long does it really take to make a realistic AI video ad? Kipp and Kieran dive into a step-by-step tutorial for creating high-quality, believable AI-powered videos, even if you're not a video expert. Learn more on how to develop a creative concept that AI tools can't replace, the essential workflow for using Veo 3.1 and Nano Banana Pro, and why reference images are the secret to seamless video scenes. This episode breaks down the process and tips to help you master AI video creation faster and smarter. Mentions Veo 3.1 https://deepmind.google/models/veo/ Nano Banana Pro https://gemini.google/overview/image-generation/ ElevenLabs https://elevenlabs.io/ Figma https://www.figma.com/ CapCut https://www.capcut.com/ Get our guide to build your own Custom GPT: https://clickhubspot.com/customgpt We're creating our next round of content and want to ensure it tackles the challenges you're facing at work or in your business. To understand your biggest challenges we've put together a survey and we'd love to hear from you! https://bit.ly/matg-research Resource [Free] Steal our favorite AI Prompts featured on the show! Grab them here: https://clickhubspot.com/aip We're on Social Media! Follow us for everyday marketing wisdom straight to your feed YouTube: https://www.youtube.com/channel/UCGtXqPiNV8YC0GMUzY-EUFg Twitter: https://twitter.com/matgpod TikTok: https://www.tiktok.com/@matgpod Join our community https://landing.connect.com/matg Thank you for tuning into Marketing Against The Grain! Don't forget to hit subscribe and follow us on Apple Podcasts (so you never miss an episode)! https://podcasts.apple.com/us/podcast/marketing-against-the-grain/id1616700934 If you love this show, please leave us a 5-Star Review https://link.chtbl.com/h9_sjBKH and share your favorite episodes with friends. We really appreciate your support. Host Links: Kipp Bodnar, https://twitter.com/kippbodnar Kieran Flanagan, https://twitter.com/searchbrat ‘Marketing Against The Grain' is a HubSpot Original Podcast // Brought to you by Hubspot Media // Produced by Darren Clarke.
7 AI Tools you'll need in 2026 (Free Guide, Prompts, Workflows): https://clickhubspot.com/ekv Ep. 389 What are the five AI skills every marketer needs to win in 2026? Kipp and Kieran dive into the top five AI launches and must-have skills transforming marketing in 2026, breaking down what matters amid an overwhelming wave of new technology. Learn more on using Gemini 3 for content remixing and competitive intelligence, mastering next-level image and video creation with cutting-edge models like Nano Banana Pro and Veo 3.1, and the critical importance of automation, agentic workflows, and vibe coding to unlock 10x marketing scale. Mentions Gemini 3 https://gemini.google.com/app Claude Opus 4.5 http://anthropic.com/news/claude-opus-4-5 Nano Banana Pro https://gemini.google/overview/image-generation/ Veo 3.1 https://deepmind.google/models/veo/ Sora 2 https://openart.ai/video/i2v/sora-v2 Get our guide to build your own Custom GPT: https://clickhubspot.com/customgpt We're creating our next round of content and want to ensure it tackles the challenges you're facing at work or in your business. To understand your biggest challenges we've put together a survey and we'd love to hear from you! https://bit.ly/matg-research Resource [Free] Steal our favorite AI Prompts featured on the show! Grab them here: https://clickhubspot.com/aip We're on Social Media! Follow us for everyday marketing wisdom straight to your feed YouTube: https://www.youtube.com/channel/UCGtXqPiNV8YC0GMUzY-EUFg Twitter: https://twitter.com/matgpod TikTok: https://www.tiktok.com/@matgpod Join our community https://landing.connect.com/matg Thank you for tuning into Marketing Against The Grain! Don't forget to hit subscribe and follow us on Apple Podcasts (so you never miss an episode)! https://podcasts.apple.com/us/podcast/marketing-against-the-grain/id1616700934 If you love this show, please leave us a 5-Star Review https://link.chtbl.com/h9_sjBKH and share your favorite episodes with friends. We really appreciate your support. Host Links: Kipp Bodnar, https://twitter.com/kippbodnar Kieran Flanagan, https://twitter.com/searchbrat ‘Marketing Against The Grain' is a HubSpot Original Podcast // Brought to you by Hubspot Media // Produced by Darren Clarke.
What does 2026 hold for indie authors and the publishing industry? I give my thoughts on trends and predictions for the year ahead. In the intro, Quitting the right stuff; how to edit your author business in 2026; Is SubStack Good for Indie Authors?; Business for Authors webinars. If you'd like to join my community and support the show every month, you'll get access to my growing list of Patron videos and audio on all aspects of the author business — for the price of a black coffee (or two) a month. Join us at Patreon.com/thecreativepenn. Joanna Penn writes non-fiction for authors and is an award-winning, New York Times and USA Today bestselling thriller author as J.F. Penn. She's also an award-winning podcaster, creative entrepreneur, and international professional speaker. You can listen above or on your favorite podcast app or read the notes and links below. Here are the highlights and the full transcript is below. (1) More indie authors will sell direct through Shopify, Kickstarter, and local in-person events (2) AI-powered search will start to shift elements of book discoverability (3) The start of Agentic Commerce (4) AI-assisted audiobook narration will go mainstream (5) AI-assisted translation will start to take off beyond the early adopters (6) AI video becomes ubiquitous. ‘Live selling' becomes the next trend in social sales. (7) AI will create, run, and optimise ads without the need for human intervention (8) 1000 True Fans becomes more important than ever You can find all my books as J.F. Penn and Joanna Penn on your favourite online store in all the usual formats, or order from your local library or bookstore. You can also buy direct from me at CreativePennBooks.com and JFPennBooks.com. I'm not really active on social media, but you can always see my photos at Instagram @jfpennauthor. 2026 Trends and Predictions for Indie Authors and Book Publishing (1) More indie authors will sell direct through Shopify, Kickstarter, and local in-person events — and more companies like BookVault will offer even more beautiful physical books and products to support this. This trend will not be a surprise to most of you! Selling direct has been a trend for the last few years, but in 2026, it will continue to grow as a way that independent authors become even more independent. The recent Written Word Media survey from Dec 2025 noted that 30% of authors surveyed are selling direct already and 30% say they plan to start in 2026. Among authors earning over $10,000 per month, roughly half sell direct. In my opinion, selling direct is an advanced author strategy, meaning that you have multiple books and you understand book marketing and have an email list already or some guaranteed way to reach readers. In fact, Kindlepreneur reports that 66% of authors selling direct have more than 5 books, and 46% have more than 10 books. Of course, you can start with the something small, like a table at a local event with a limited number of books for sale, but if you want to consistently sell direct for years to come, you need to consider all the business aspects. Selling direct is not a silver bullet. It's much harder work to sell direct than it is to just upload an ebook to Amazon, whether you choose a Kickstarter campaign, or Shopify/Payhip or other online stores, or regular in-person sales at events/conferences/fairs. You need a business mindset and business practices, for example, you need to pay upfront for setup as well as ongoing management, and bulk printing in some cases. You need to manage taxes and cashflow. You need to be a lot more proactive about marketing, as you won't sell anything if you don't bring readers to your books/products. But selling direct also brings advantages. It sets you apart from the bulk of digital only authors who still only upload ebooks to Amazon, or maybe add a print on demand book, and in an era of AI rapid creation, that number is growing all the time. If you sell direct, you get your customer data and you can reach those customers next time, through your email list. If you don't know who bought your books and don't have a guaranteed way to reach them, you will more easily be disrupted when things change — and they always change eventually. Kindlepreneur notes that “45% of the successful direct selling authors had over 1,000 subscribers on their email lists,” with “a clear, positive correlation between email list size and monthly direct sales income — with authors having an email list of over 15,000 subscribers earning 20X more than authors with email lists under 100 subscribers.” Selling direct means faster money, sometimes the same day or the same week in many cases, or a few weeks after a campaign finishes, as with Kickstarter. And remember, you don't have to sell all your formats directly. You can keep your ebooks in KU, do whatever you like with audiobooks, and just have premium print products direct, or start with a very basic Kickstarter campaign, or a table at a local fair. Lots more tips for Shopify and Kickstarter at https://www.thecreativepenn.com/selldirectresources/ I also recommend the Novel Marketing Podcast on The Shopify Trap: Why authors keep losing money as it is a great counterpoint to my positive endorsement of selling direct on Shopify! Among other things, Thomas notes that a fixed monthly fee for a store doesn't match how most authors make money from books which is more in spikes, the complexity and hassle eats time and can cost more money if you pay for help, and it can reduce sales on Amazon and weaken your ranking. Basically, if you haven't figured out marketing direct to your store, it can hurt you.All true for some authors, for some genres, and for some people's lifestyle. But for authors who don't want to be on the hamster wheel of the Amazon algorithm and who want more diversity and control in income, as well as the incredible creative benefits of what you can do selling direct, then I would say, consider your options in 2025, even if that is trying out a low-financial-goal Kickstarter campaign, or selling some print books at a local fair. Interestingly, traditional publishers are also experimenting with direct sales. Kate Elton, the new CEO of Harper Collins notes in The Bookseller's 2026 trend article, “we are seeing global success with responsive, reader-driven publishing, subscription boxes and TikTok Shop and – crucially – developing strategies that are founded on a comprehensive understanding of the reader.” She also notes, “AI enables us to dramatically change the way we interact with and grow audiences. The opportunities are genuinely exciting – finding new ways to help readers discover books they will love, innovating in the ways we market and reach audiences, building new channels and adapting to new methods of consuming content.” (2) AI-powered search will start to shift elements of book discoverability From LinkedIn's 2026 Big Ideas: “Generative engine optimization (GEO) is set to replace search engine optimization (SEO) as the way brands get discovered in the year ahead. As consumers turn to AI chatbots, agentic workflows and answer engines, appearing prominently in generative outputs will matter more than ranking in search engines.” Google has been rolling out AI Mode with its AI Overviews and is beginning to push it within Google.com itself in some countries, which means the start of a fundamental change in how people discover content online. I first posted about GEO (Generative Engine Optimisation) and AEO (Answer Engine Optimisation) in 2023, and it's going to change how readers find books. For years, we've talked about the long tail of search. Now, with AI-powered search, that tail is getting even longer and more nuanced. AI can understand complex, conversational queries that traditional search engines struggled with. Someone might ask, “What's a good thriller set in a small town with a female protagonist who's a journalist investigating a cold case?” and get highly specific recommendations. This means your book metadata, your website content, and your online presence need to be more detailed and conversational. AI search engines understand context in ways that go far beyond simple keywords. The authors who win in this new landscape will be those who create rich, authentic content about their books and themselves, not just promotional copy. As economist Tyler Cowen has said, “Consider the AIs as part of your audience. Because they are already reading your words and listening to your voice.” We're in the ‘organic' traffic phase right now, where these AI engines are surfacing content for ‘free,' but paid ads are inevitably on the way, and even rumoured to be coming this year to ChatGPT. By the end of 2026, I expect some authors and publishers to be paying for AI traffic, rather than blocking and protesting them. For now, I recommend checking that your author name/s and your books are surfaced when you search on ChatGPT.com as well as Google.com AI Mode (powered by Gemini). You want to make sure your work comes up in some way. I found that Joanna Penn and J.F. Penn searches brought up my Shopify stores, my website, podcast, Instagram, LinkedIn, and even my Patreon page, but did not bring up links to Amazon. If you only have an author presence on Amazon, does it appear in AI search at all? Do you need to improve anything about what the AI search brings up? Traditional publishers are also looking at this, with PublishersWeekly doing webinars on various aspects of AI in early 2026, including sessions on GEO and how book sales are changing, AI agents, and book marketing. In a 2026 predictions article on The Bookseller, the CEO of Bloomsbury Publishing noted, “The boundaries of artificial intelligence will become clearer, enabling publishers to harness its benefits while seeking to safeguard the intellectual property rights of authors, illustrators and publishers.” “AI will be deeply embedded in our workflows, automating tasks such as metadata tagging, freeing teams to focus on creativity and strategy. Challenges will persist. Generative AI threatens traditional web traffic and ad revenue models, making metadata optimisation and SEO critical for visibility as we adjust to this new reality online.” (3) The start of Agentic Commerce AI researches what you want to buy and may even buy on your behalf. Plus, I predict that Amazon does a commerce deal with OpenAI for shopping within ChatGPT by the end of 2026. In September 2025, ChatGPT launched Instant Checkout and the Agentic Commerce Protocol, which will enable bots to buy on websites in the background if authorised by the human with the credit card. VISA is getting on board with this, so is PayPal, with no doubt more payment options to come. In the USA, ChatGPT Plus, Pro, and Free users can now buy directly from US Etsy sellers inside the chat interface, with over a million Shopify merchants coming soon. Shopify and OpenAI have also announced a partnership to bring commerce to ChatGPT. I am insanely excited about this as it could represent the first time we have been able to more easily find and surface books in a much more nuanced way than the 7 keywords and 3 categories we have relied on for so long! I've been using ChatGPT for at least the last year to find fiction and non-fiction books as I find the Amazon interface is ‘polluted' by ads. I've discovered fascinating books from authors I've never heard of, most in very long tail areas. For example, Slashed Beauties by A. Rushby, recommended by ChatGPT as I am interested in medical anatomy and anatomical Venuses, and The Macabre by Kosoko Jackson, recommended as I like art history and the supernatural. I don't think I would have found either of these within a nuanced discussion with ChatGPT. Even without these direct purchase integrations, ChatGPT now has Shopping Research, which I have found links directly to my Shopify store when I search for my books specifically. Walmart has partnered with OpenAI to create AI-first shopping experiences, and you have to wonder what Amazon might be doing? In Nov 2025, Amazon signed a “strategic partnership” with OpenAI, and even though it's focused on the technical side of AI, those two companies in a room together might also be working on other plans … I'm calling it for 2026. I think Amazon will sign a commerce agreement with OpenAI sometime before the end of the year. This will enable at least recommendation and shopping links into Amazon stores (presumably using an OpenAI affiliate link), or perhaps even Instant Checkout with ChatGPT for Amazon. It will also enable a new marketing angle, especially if paid ads arrive in ChatGPT, perhaps even integrating with Amazon Ads in some way as part of any possible agreement, since ads are such a good revenue stream for Amazon anyway. The line between discovery, engagement, and purchase is collapsing. Someone could be having a conversation with an AI about what to read next, and within that same conversation, purchase a bookwithout ever leaving the chat interface. This already happens within TikTok and social commerce clearly works for many authors. It's possible that the next development for book discoverability and sales might be within AI chats. This will likely stratify the already fragmented book eco-system even more. Some readers will continue to live only within the Amazon ecosystem and (maybe) use their Rufus chatbot to buy, and others will be much wider in their exploration of how to find and discover books (and other products and services). If you haven't tried it yet, try ChatGPT.com Shopping Research for a book. You can do this on the free tier. Use the drop down in the main chat box and select Shopping Research. It doesn't have to be for your book. It can be any book or product, for example, our microwave died just before Christmas so I used it to find a new one. But do a really nuanced search with multiple requirements. Go far beyond what you would search for on Amazon. In the results, notice that (at the time of writing) it does not generally link to Amazon, but to independent sites and stores. As above, I think this will change by the end of 2026, as some kind of commerce deal with Amazon seems inevitable. (4) AI-assisted audiobook narration will go mainstream I've been talking about AI narration of audiobooks since 2019, and over the years, I've tried various different options. In 2025, the technology reached a level of emotional nuance that made it much easier to create satisfying fiction audio as well as non-fiction. It also super-charges accessibility, making audio available in more languages and more accents than ever before. Of course, human narration remains the gold standard, but the cost makes it prohibitive for many authors, and indeed many small traditional publishers, for all books. If it costs $2000 – $10,000 to create an audiobook, you have to sell a lot to make a profit, and the dominance of subscription models have made it harder to recoup the costs. Famous narrators and voice artists who have an audience may still be worth investing in, as well as premium production, but require an even higher upfront cost and therefore higher sales and streams in return. AI voice/audio models are continuing to improve, and even as this goes out, there are rumours on TechCrunch that OpenAI's new device, designed by Jony Ive who designed the iPhone, will be audio first and OpenAI are improving their voice models even more in preparation for that launch. In 2026, I think AI-narrated audio will go mainstream with far-reaching adoption across publishing and the indie author world in many different languages and accents. This will mean a further stratification of audiobooks, with high quality, high production, high cost human narrated audio for a small percentage of books, and then mass market, affordable AI-narrated audio for the rest. AI-narrated audiobooks will make audio ubiquitous, and just as (almost) every print book has an ebook format, in 2026, they will also have an audio format. I straddle both these worlds, as I am still a human audiobook narrator for my own work. I human-narrated Successful Self-Publishing Fourth Edition (free audiobook) and The Buried and the Drowned, my short story collection. I also use AI narration for some books. ElevenLabs remains my preferred service and in 2025, I used my J.F. Penn voice clone for Death Valley and also Blood Vintage, while using a male voice for Catacomb. I clearly label my AI-narration in the sales description and also on the cover, which I think is important, although it is not always required by the various services. You can distribute ElevenLabs narrated audiobooks on Spotify, Kobo Writing Life, YouTube, ElevenReader, and of course your own store if you use Shopify with Bookfunnel. There are many other services springing up all the time, so make sure you check the rights you have over the finished audio, as well as where you can sell and distribute the final files. If they are just using ElevenLabs models in the back-end, then why not just do that directly? (Most services will be using someone's model in the back-end, since most companies do not train their own models.) Of course, you can use Amazon's own narration. While Amazon originally launched Audible audiobooks with Virtual Voice (AVV) in November 2023, it was rolled out to more authors and territories in 2025. If your book is eligible, the option to create an audiobook will appear on your KDP dashboard. With just a few clicks, you can create an audiobook from a range of voices and accents, and publish it on Amazon and Audible. However, the files are not yours. They are exclusive to Amazon and you cannot use them on other platforms or sell them direct yourself. But they are also free, so of course, many authors, especially those in KU, will use this option. I have done some for my mum's sweet romance books as Penny Appleton and I will likely use them for my books in translation when the option becomes available. Traditional publishers are experimenting with AI-assisted audiobook narration as well. MacMillan is selling digital audiobooks read by AI directly on their store. PublishersWeekly reports that PRH Audio “has experimented with artificial voice in specific instances, such as entrepreneur Ely Callaway's posthumous memoir The Unconquerable Game,” when an “authorized voice replica” was created for the audiobook. The article also notes that PRH Audio “embrace artificial intelligence across business operations—my entire department [PRH Audio] is using AI for business applications.” And while indie authors can't use AI voices on ACX right now, Audible have over 100 voices available to selected publishing partnerships, as reported by The Guardian with “two options for publishers wishing to make use of the technology: “Audible-managed” production, or “self-service” whereby publishers produce their own audiobooks with the help of Audible's AI technology.” In 2026, it's likely that more traditional publishers — as well as indie authors — will get their backlist into audio with AI narration. (5) AI-assisted translation will start to take off beyond the early adopters Over the years, I've done translation deals with traditional publishers in different languages (German, French, Spanish, Korean, Italian) for some fiction and non-fiction books. But of course, to get these kinds of deals, you have to be proactive about pitching, or work with an agent for foreign rights only, and those are few and far between! There are also lots of languages and territories worldwide, and most deals are for the bigger markets, leaving a LOT of blue water for books in translation, even if you have licensed some of the bigger markets. I did my first partially AI-translated books in 2019 when I used Deepl.com for the first draft and then worked with a German editor to do 3 non-fiction books in German. While the first draft was cheap, the editing was pretty expensive, so I stopped after only doing a couple. I have made the money back now, but it took years. In 2025, AI Translation began to take off with ScribeShadow, GlobeScribe.ai, and more recently, in November 2025, Kindle Translate boosting the number of translated books available. Kindle Translate is (currently) only available to US authors for English into Spanish and also German into English, but in 2026, this will likely roll out to more languages and more authors, making it easier than ever to produce translations for free. Of course, once again, the gold standard is human translation, or at least human-edited translations, but the cost is prohibitive even just for proof-reading, and if there is a cheap or even free option, like Kindle Translate, then of course, authors are going to try it. If the translation gets bad reviews, they can just un-publish. There are many anecdotal stories of indie success in 2025 with AI-translated genre fiction sales (in series) in under-served markets like Italian, French, and Spanish, as well as more mainstream adoption in German. I was around in the Kindle gold-rush days of 2009-2012 and the AI-translation energy right now feels like that. There are hardly any Kindle ebooks in many of these languages compared to how many there are in English, so inevitably, the rush is on to fill the void, especially in genres that are under-served by traditional publishers in those markets. Yes, some of these AI translated books will be ‘AI-slop,' but readers are not stupid. Those books will get bad reviews and thus will sink to the bottom of the store, never to be seen again. The AI translation models are also improving rapidly, and Amazon's Kindle Translate may improve faster than most, for books specifically, since they will be able to get feedback in terms of page reads. Amazon is also a major investor in Anthropic, which makes Claude.ai, widely considered the best quality for creative writing and translation, so it's likely that is used somewhere in the mix. Some traditional publishers are also experimenting with AI-assisted translation, with Harlequin France reportedly using AI translation and human proofreaders, as reported by the European Council of Literary Translators' Associations in December 2025. Academic publisher Taylor and Francis is also using AI for book translation, noting: “Following a program of rigorous testing, Taylor & Francis has announced plans to use AI translation tools to publish books that would otherwise be unavailable to English-language readers, bringing the latest knowledge to a vastly expanded readership.” “Until now, the time and resources required to translate books has meant that the majority remained accessible only to those who could read them in the original language. Books that were translated often only became available after a significant delay. Today, with the development of sophisticated AI translation tools, it has become possible to make these important texts available to a broad readership at speed, without compromising on accuracy.” (6) AI video becomes ubiquitous. ‘Live selling' becomes the next trend in social sales. In 2025, short form AI-generated video became very high quality. OpenAI released Sora 2, and YouTube announced new Shorts creation tools with Veo 3, which you can also use directly within Gemini. There are tons of different AI video apps now, including those within the social media sites themselves. There is more video than ever and it's much easier to create. I am not a fan of short form video! I don't make it and I don't consume it, but I do love making book trailers for my Kickstarter campaigns and for adding to my book pages and using on social media. I made a trailer for The Buried and the Drowned using Midjourney for images and then animation of those images, and Canva to put them together along with ElevenLabs to generate the music. But despite the AI tools getting so much easier to use, you still have to prompt them with exactly what you want. I can't just upload my book and say, “Make a book trailer,” or “Make a short film.” This may change with generative video ads, which are likely to become more common in 2026, as video turns specifically commercial. Video ads may even be generated specifically for the user, with an audience of one, maybe even holding your book in their hands (using something like Cameos on Sora), in the same way that some AI-powered clothing stores do virtual try-ons. This might also up-end the way we discover and buy things, as the AI for eCommerce and Amazon Sellers newsletter says about OpenAI's Sora app, “OpenAI isn't just trying to build a TikTok competitor. They're building a complete reimagining of how we discover and buy things …” “The combination of ChatGPT's research capabilities and Sora's potential for emotional manipulation—I mean, “engagement”—could create something we've never seen before: an AI ecosystem that might eventually guide you through every type of purchase, from the most considered to the most impulsive.” In 2026, there will be A LOT more AI-generated video, but that also leads to the human trend of more live video. While you can use an AI avatar that looks and sounds like you using tools like HeyGen or Synthesia, live video has all the imperfect human elements that make it stand-out, plus the scarcity element which leads to the purchase decision within a countdown period. Live video is nothing new in terms of brand building and content in general, but it seems that live events primarily for direct sales might be a thing in 2026. Kim Kardashian hosted Kimsmas Live in December 2025 with a 45 minute live shopping event with special guests, described as entertainment but designed to be a sales extravaganza. Indie authors are doing a similar thing on TikTok with their books, so this is a trend to watch in 2026, especially if you feel that live selling might fit with your personality and author business goals. It's certainly not for everyone, but I suspect it will suit a different kind of creator to those who prefer ‘no face' video, or no video at all! On other aspects of the human side of social media, Adam Mosseri the CEO of Instagram put a post on Threads called Authenticity after Abundance. He said, “Everything that made creators matter—the ability to be real, to connect, to have a voice that couldn't be faked—is now suddenly accessible to anyone with the right tools.” “Deepfakes are getting better and better. AI is generating photographs and videos indistinguishable from captured media. The feeds are starting to fill up with synthetic everything. And in that world, here's what I think happens.Creators matter more.” It's a long article so just to pick a few things from it: “We like to talk about “AI slop,” but there is a lot of amazing AI content … we are going to start to see more and more realistic AI content.” I've talked to my Patreon Community about this ‘tsunami of excellence' as these tools are just getting better and better and the word ‘slop' can also be applied to purely human output, too. If you think that AI content is ‘worse' than wholly human content, in 2026, you are wrong. It is now very very good, especially in the hands of people who can drive the AI tools. Back to Adam's post: “Authenticity is fast becoming a scarce resource, …The creators who succeed will be those who figure out how to maintain their authenticity [even when it can be simulated] …” “The bar is going to shift from “can you create?” to “can you make something that only you could create?” He talks about how the personal content on Instagram now is: “unpolished; it's blurry photos and shaky videos of people's daily experiences … flattering imagery is cheap to produce and boring to consume. People want content that feels real… Savvy creators are going to lean into explicitly unproduced and unflattering images of themselves. In a world where everything can be perfected, imperfection becomes a signal. Rawness isn't just aesthetic preference anymore—it's proof. It's defensive. A way of saying: this is real because it's imperfect.” While I partially love this, and I really hope it's true, as in I hope we don't need to look good for the camera anymore I would also challenge Adam on this, because pretty much every woman I know on social media has been sent sexual messages, and/or told they are ugly and/or fat when posting anything unflattering. I've certainly had both even for the same content, but I don't expect Adam has been the target for such posting! But I get his point. He goes on:“Labeling content as authentic or AI-generated is only part of the solution though. We, as an industry, are going to need to surface much more context about not only the media on our platforms, but the accounts that are sharing it in order for people to be able to make informed decisions about what to believe. Where is the account? When was it created? What else have they posted?” This is exactly what I've been saying for a while under my double down on being human focus. I use my Instagram @jfpennauthor as evidence of humanity, not as a sales channel. You can do both of course, but increasingly, you need to make sure your accounts at places have longevity and trust, even by the platforms themselves. Adam finishes: “In a world of infinite abundance and infinite doubt, the creators who can maintain trust and signal authenticity—by being real, transparent, and consistent—will stand out.” For other marketing trends for 2026, I recommend publicist Kathleen Schmidt's SubStack which is mostly focused on traditional publishing but still interesting for indies. In her 2026 article, she notes: “We have reached a social media saturation point where going viral can be meaningless and should not be the goal; authenticity and creativity should. She also says, “In-person events are important again,” and, “Social media marketing takes a nosedive… we have reached a saturation point … What publishers must figure out is how to make their social media campaigns stand out. If they remain somewhat uninspired, the money spent on social ads won't convert into book sales.” I think this is part of the rise of live selling as above, which can stand out above more ‘produced' videos. Kathleen also talks about AI usage. “AI can help lighten the burden of publicity and marketing.” “A lot of AI tools are coming to market to lessen the load: they can write pitches, create media lists for you, send pitches for you, and more. I know the industry is grappling with all things AI, but some of these tools are huge time savers and may help a book more than hurt it.” On that note … (7) AI will create, run, and optimise ads without the need for human intervention Many authors will be very happy about this as marketing is often the bane of our author business lives! As I noted in my 2026 goals, I would love to outsource more marketing tasks to AI. I want an “AI book marketing assistant” where I can upload a book and specify a budget and say, ‘Go market this,' then the AI will action the marketing, without me having to cobble together workflows between systems. Of course, it will present plans for me to approve but it will do the work itself on the various platforms and monitor and optimize things for me. I really hope 2026 is the year this becomes possible, because we are on the edge of it already in some areas. Amazon Ads launched a new agentic AI tool in September 2025 that creates professional-quality ads. I've also been working with Claude in Chrome browser to help me analyse my Amazon Ad data and suggest which keywords/products to turn off and what to put more budget into. I'll do a Patreon video on that soon. Meta announced it will enable AI ad creation by the end of 2026 for Facebook and Instagram. For authors who find ad creation overwhelming or time-consuming, this could be a game-changer. Of course, you will still need a budget! (8) 1000 True Fans becomes more important than ever Lots of authors and publishers are moaning about the difficulty of reaching readers in an era of ‘AI slop' but there is no shortage of excellent content created by humans, or humans using AI tools. As ever, our competition is less about other authors, or even authors using AI-assisted creation, we're competing against everything else that jostles for people's attention, and the volume of that is also growing exponentially. I've never been a fan of rapid release, and have said for years that you can't keep up with the pace of the machines. So play a different game. As Kevin Kelly wrote in 2008, If you have 1000 true fans, (also known as super fans), “you can make a living — if you are content to make a living but not a fortune.” [Kevin Kelly was on this show in 2023 talking about Excellent Advice for Living.] Many authors and the publishing industry are stuck in the old model of aiming to sell huge volumes of books at a low profit margin to a massive number of readers, many of them releasing ever faster to try and keep the algorithms moving. But the maths can work for the smaller audience of more invested readers and fans. If you only make $2 profit on an ebook, you need to sell 500 ebooks to make $1000, and then do it again next month. Or you can have a small community like my patreon.com/thecreativepenn where people pay $2 (or more) a month, so even a small revenue per person results in a better outcome over the year, as it is consistent monthly income with no advertising. But what if you could make $20 profit per book? That is entirely possible if you're producing high quality hardbacks on Kickstarter, or bundle deals of audiobooks, or whole series of ebooks. You would only need to sell to 50 people to make $1000. What about $100 profit per sale, which you can do with a small course or live event? You only need 10 people to make $1000, and this in-person focus also amplifies trust and fosters human connection. I've found the intimacy of my live Patreon Office Hours and also my webinars have been rewarding personally, but also financially, and are far more memorable — and potentially transformative — than a pre-recorded video or even another book. From the LinkedIn 2026 Big Ideas article: “In an AI-optimized world, intentional human connection will become the ultimate luxury.” The 1000 True Fans model is about serving a smaller, more personal audience with higher value products (and maybe services if that's your thing). As ever, its about niche and where you fit in the long long long long long tail. It's also about trust. Because there is definitely a shortage of that in so many areas, and as Adam Mosseri of Instagram has said, trust will be increasingly important. Trust takes time to build, but if you focus on serving your audience consistently, and delivering a high quality, and being authentic, this emerges as part of being human. In an echo of what happened when online commerce first took off, we are back to talking about trust. Back in 2010, I read Trust Agents: by Julien Smith and Chris Brogan, which clearly needs a comeback. There was a 10th anniversary edition published in 2020, so that's worth a read/listen. Chris Brogan was also on this show in 2017 when we talked about finding and serving your niche for the long term. That interview is still relevant, here's a quick excerpt, where I have (lightly edited) his response to my question on this topic back in 2017: Jo: The principle of know, like, and trust, why is that still important or perhaps even more important these days? Chris: There are a few things that at play there, Joanna. One is that the same tools that make it so easy for any of us to start and run a business also allow certain elements to decide whether or not they want to do something dubious. And with all new technologies that come, you know, there's nothing unique about these new technologies. In the 1800s, anyone could put anything in a bottle and sell it to you and say, this is gonna cure everything. Cancer — gone. And the bottle could have nothing in. You know, it could be Kool-Aid. And so, the idea of trying to understand what's behind the business though, one beautiful thing that's come is that we can see in much more dimensions who we're dealing with. We can understand better who's the face behind the brand. I really want people to try their best to be a lot clearer on what they stand for or what they say. And I don't really mean a tagline. I mean, humans don't really talk like that. They don't throw some sentence out as often as they can that you remember them for that phrase. But I would say that, we have so many media available to us — the plural of mediums — where we can be more of ourselves. And I think that there's a great opportunity to share the ‘you' behind the scenes, and some people get immediately terrified about this, ‘Ah, the last thing I want is for people to know more about me,' but I think we have such an opportunity. We have such an opportunity to voice our thoughts on something, to talk about the story that goes behind the product. We were all raised on overly produced material, but I think we don't want that anymore. We really want clarity, brevity, simplicity. We want the ability for what we feel is connection and then access. And so I think it's vital that we connect and show people our accessibility, not so that they can pester us with strange questions, but more so that you can say, this person stands with their product and their service and this person believes these things, and I feel something when I hear them and I wanna be part of that.” That's from Chris Brogan's interview here in 2017, and he is still blogging and speaking at writing at ChrisBrogan.com and I'm going to re-listen to the audiobook of Trust Agents again myself as I think it's more relevant than ever. The original quote comes from Bob Burg in his 1994 book, Endless Referrals, “All things being equal, people will do business with, and refer business to, those people they know, like and trust.” That still applies, and absolutely fits with the 1000 True Fans model of aiming to serve a smaller audience. As Kevin Kelly says in 1000 True Fans, “Instead of trying to reach the narrow and unlikely peaks of platinum bestseller hits, blockbusters, and celebrity status, you can aim for direct connection with a thousand true fans.” “On your way, no matter how many fans you actually succeed in gaining, you'll be surrounded not by faddish infatuation, but by genuine and true appreciation. It's a much saner destiny to hope for. And you are much more likely to actually arrive there.” In 2026, I hope that more authors (including me!) let go of ego goals and vanity metrics like ranking, gross sales (income before you take away costs), subscribers, followers, and likes, and consider important business numbers like profit (which is the money you have after costs like marketing are taken out), as well as number of true fans — and also lifestyle elements like number of weekends off, or days spent enjoying life and not just working! OK, that's my list of trends and predictions for 2026. Let me know what you think in the comments. Do you agree? Am I wrong? What have I missed? The post 2026 Trends And Predictions For Indie Authors And The Book Publishing Industry with Joanna Penn first appeared on The Creative Penn.
Did AI end up being a political force this year?