Nate B Jones

Anthropic’s Mythos Just Beat OpenAI’s GPT-5.5 At Real Hacking

🇬🇧 EN🇪🇸 ES
24:18 min youtube 2026 Semana 20 🇪🇸 ES
Transcripción completa
[00:00] A guy on X posted this week that he just recovered five Bitcoin from a wallet he'd locked himself out of 11 years ago. It's worth about $400,000 and he thought it was gone forever. The story is that he changed the password back in college, got intoxicated, and then forgot the new password. Years of brute force crackers and password dictionaries had no effect. What finally got it back wasn't a better password cracker, it was Claude. He uploaded files from his old college hard drive and let Claude sort through more than a decade of forgotten folders.
[00:30] Claude found an older wallet.dat file from before the password change, lined it up with a mnemonic recovery phrase he still had, and the wallet opened. He didn't hack anything. He just had an AI do what a patient research assistant would do for as long as it took. I think that's really illustrative as a story of where we are right now with AI. The model launches are still happening, but the more interesting stuff is quieter and much more specific. Agents are starting to do real work on real artifacts inside real companies and with
[01:00] real people, and it's beginning to change real life, real decisions for people building products. Skipping past the Bitcoin one, there are five stories from this week that are worth paying attention to, and each one of them changes a different decision for you. Notion launched a real developer platform for agents, Anthropic tightened Claude's usage limits again. There's new data suggesting Anthropic has crossed a business adoption line a lot of people assumed OpenAI owned. Uh Mythos and GPT-5.5 made it clear that AI cyber capability is moving faster than most
[01:31] people are really ready to process. And AWS gave agents managed cloud desktops, which sounds kind of boring until you remember how much company work still lives in software without an API. So, those are the five stories we're going to get into. Let's jump into Notion first. Notion gave agents a front door into the workspace. On May 13th, Notion launched their developer platform. And I know that seems like it's just for engineers, but actually all of us are going to benefit from this, and I can explain why. So, there's a Notion command line interface, and that's built for developers, people comfortable with
[02:01] a command line, and people who use coding agents to access tools. But, then on top of that, there are what Notion calls workers, hosted functions that run on Notion's own infrastructure. For example, there's database sync, so you can pull data from Salesforce, from Stripe, from GitHub, from Zendesk, from Postgres, or any API into a Notion database and keep that data fresh with an automatic worker job. There are also webhooks that trigger Notion from outside systems, uh custom agent tools, and an external agents API that lets you
[02:31] bring agents like Claude and Codex or others into your own internal Notion workspace as participants. So, Notion is not just adding an AI button, nor are they leaning on their very popular Notion AI tooling that they already have inside the system, which I've covered before. They're actually trying to go farther and make the entire workspace programmable. That matters because so much company work does not start in a formal enterprise system. It starts in a Notion doc, in a project database, a customer notes page, an operating
[03:02] checklist, or even in a a rogue project, right? Like a lightweight CRM somebody built because Salesforce was too heavy and a spreadsheet was too brittle. Those are exactly the awkward places and corners where agents need a lot of context, and until now you didn't have a good way to get them in there. Before this, your options were pretty awkward. You could use Notion as a human workspace and run some agents that Notion built inside that workspace, and then you had to run a bunch of agent work elsewhere for other stuff. You could build some brittle glue around the Notion API and try and make it work, or
[03:34] you could have an agent read a Notion page and summarize it, which is useful but not enough for serious work. The new platform is much more foundational to the agentic world we're all entering. A coding agent can write a worker, the worker can sync data to a Notion database, a webhook can trigger that worker, an agent can use the sync context, and a human can review the results all in the same workspace where the team already works. Take customer onboarding. A deal closes in Salesforce, a web hook fires a Notion worker, the worker spins up an
[04:04] onboarding workspace, pulls in plan data, account notes, success criteria, milestones, and support history. An agent drafts the kickoff plan, the CSM reviews it in Notion before anything goes to the customer, and none of that is speculation by the way. That is all like foundational components that this release was built for that form a coherent workflow. So, if your company already lives in Notion, this is a big week. You can pick a database that already matters to you and dig in with projects, customers, candidates, support issues, whatever, and you can ask yourself "Do I have outside data I need
[04:35] to bring in here? Is there an event that should trigger work that I wish I could have done before? What What should the agent draft, update, and check? Where does a human approve the final step? Do I want other agents in this space like Claude or Codex?" All of that is now on the table. So, the Notion headline is not Notion launched AI. That already happened. I talked about it. It's instead that Notion is expanding on their AI launch to become the workbench where humans and agents share context, or at least that's their goal. All right, story number two. Claude limits got tighter for developers because agent
[05:06] usage is breaking the subscription model. And this is really a story about the perils of success because Claude was the product that kicked off the agent revolution back in December, and ever since then we've seen this enormous runaway ramp of AI usage driven by agents, and now Claude is running out of compute, and the Anthropic team is having to figure out what they're going to do about it. So, let's get into it. Axios reported this week that Anthropic is moving some outside agent tool usage behind its own credit meter. Sam Altman responded to that by offering new
[05:36] business customers 2 months of free Codex. This is not just a promo fight between two heavyweights. It's the market figuring out that all you can eat AI means something very different when the user isn't a person typing into a chatbot, right? For builders, the implication is really clear. Usage limits are now a product behavior you have to pay attention to. If your AI workflow depends on Claude code, Codex, Cursor, or any agent that runs for a lengthy period of time, the limit that they're talking about is not just an accounting detail or a billing detail. It's actually part of the user
[06:06] experience. What happens when the agent hits a billing cap halfway through a task? Does the work pause? Does it resume? Does it switch models? Does it bill you more? Does it lose context? Does the user even understand what happened? Does your team know the cost per completed task? Most teams don't have good answers to this stuff yet. They're still thinking in seats. They're thinking in subscriptions. One of the things we need to understand is that framing work as agentic isn't enough. The second story is about Claude. Now, Claude has been in hot water because of the runaway success of
[06:38] agentic workflows since January, since December, which ironically, Claude itself kicked off, right? Do you remember back in December when Claude code started to become really good and everyone started switching to Claude? That's what kicked off the last 6 months. It's been incredible. The Anthropic team is out of computers. In early April, mid-April 2026, they began to clamp down on runaway agentic flows, and they started very unpopularly by clamping down on open Claude. Yes, pun intended. When they did that, when they cut off open Claude and other
[07:08] third-parties from using the Claude subscriptions that people are paying for, what happened was most developers who'd been used to consuming thousands of dollars in tokens and paying hundreds kind of got upset about that. That was a deal that worked really well for them, and that was not a deal that worked well for Anthropic. And so, Anthropic initially responded that by saying essentially, "No more third-party usage of your personal subscription costs. That's far too financially unsustainable for us. We we cannot do that. Instead,
[07:39] please use our API and pay by the token." And developers, many of whom were building open claw, building other side projects, couldn't necessarily afford to do that. And so, it killed a lot of projects. It drove a lot of people over to open AI. And open AI welcomed them with outstretched arms. There was a lot of work that the open AI team did to make sure that open claw in particular was friendly to open AI. And it was easy to use your open AI subscription with open claw. So, they basically came back and said, "Hey, this is a user acquisition opportunity." Well, now, Anthropic kind of wants to
[08:10] come back and make a gesture toward the open claw community and toward agents that individuals are running as side projects. And so, what they've decided to do is instead of saying, as they did in April, "You may not use your 20 or $200 a month subscription for any third-party agent usage." They're now saying, "Well, you can, but there's a rate limit on it. And the rate limit is something that expires every month. It's use it or lose it. And then after that, you have to pay the buy-the-token API billing cost." That has not been popular. One of the things about PR with
[08:42] developers is that the clearer, simpler, more consistent message wins. And one of the underlying challenges for Anthropic right now is that they were the ones that had the clearer, simpler message for developers for a long, long time. Part of why they broke through in December and January is that they were so clear and so simple about Claude code being a delightful experience. But now, because of its success, they can't afford to be that clear and simple with developers anymore. And it is costing them a lot of goodwill with the
[09:12] developer community. Now, I don't want to overstate that. I still know a ton of developers who use Claude and Claude code. Let's not pretend that they don't. That is not what I am saying. There's still massive adoption, and you see that in the revenue numbers for Anthropic. But, it is painful to try to explain to developers that their usage has been financially unsustainable after they have gotten used to it. And that is the battle that Anthropic is is fighting right now. The third story is a revenue one. For most of the Chat GPT era,
[09:42] OpenAI has been the gorilla in the room when it comes to revenue. That has all changed in the last 6 months, and now it is fair to say that Anthropic and OpenAI are neck and neck by most revenue terms. It's really hard to compare them directly because we tend to see trailing indicators as far as the news that hits the markets, the news that leaks out of these two companies. The baseline assumption is that both companies are getting close to or a little bit over $30 in annualized revenue, and they're both neck and neck. One hard fact that doesn't come from either company is from
[10:13] Ramp, which does know how companies are spending their money because that's their business. They handle cards and payments for companies. And what they have called out is that for the first time Anthropic has more verified business customers than OpenAI. And so that is a metric that indicates how close this race is and how quickly Anthropic has gained ground in the last few months. I think one of the larger stories here is that revenue is a little bit of a leading edge indicator of health. I know we often talk about revenue in business as a trailing edge indicator, and for a lot of business modeling it is. But in this situation,
[10:44] revenue is a leading edge indicator for the strain on compute that all of these companies are going to put on the supply chain. And so what Anthropic is dealing with is the reality that they had planned for 10x growth in a year, and they're over 80x. And those aren't my words, that's actually straight from Dario Amodei, and he's saying, "I underplanned for growth, and I'm trying to find the compute to support the sort of growth we've had." It's a great problem to have, but it is a real problem, and it's a problem that Anthropic is going to have to deal with if they plan to keep that
[11:14] revenue sustainable going forward. The fourth story this week is another significant indicator that Methos is a special model when it comes to cybersecurity. This week two independent evaluations of Claude Methos preview dropped, and both of them are absolutely compelling from a cybersecurity perspective. One was from the XPODW organization and the other one was from the UK AI Security Institute. And the short version of this story is that frontier models, especially Mythos, are now good enough at serious cyber work
[11:44] that security teams need to update their assumptions. And I'm and I'm underlining that because previously, that has been the story Anthropic has told us, but it's also the story they're incentivized to tell us. And so you have to take it with a grain of salt. We're now seeing independent evaluators look at Mythos and say, "Wow, no, there is something really cool here. There's something really significant. There's something that we as security researchers need to pay close attention to here." Let's get into what they discovered. First, the tests were serious tests. The
[12:14] model has to work through an actual attack chain, looking at reconnaissance, credential theft, lateral movement through the attack surface, web app exploitation, privilege escalation, command and control persistence, infrastructure compromise, and finally full network takeover. Those are all stacked in layer of difficulty. And the AI Security Institute tested that entire attack chain with multiple models and put together a chart that compares the performance of Mythos preview, GPT 5.5, 5.5 cyber, other Claude Opus models,
[12:44] other Codex models, and older models. Mythos preview gets farther in that attack chain on the same token budget than any other model. And that matters partly because ChatGPT 5.5 is an extraordinarily good model. OpenAI positioned it as more token efficient than 5.4, and it absolutely is. And in OpenAI's own evals, 5.5 is ahead of Opus 4.7 on the cybersecurity benchmark Cyber Gym. The AI Security Institute said 5.5
[13:16] itself significantly exceeds the old cyber progress trend. It shows that it's good. So Mythos winning here is not Mythos beating a weak baseline. It's actually outrunning an extremely strong model on a task where token spend is a metric that matters. And And that's partly because in cybersecurity cost matters. If a model can find real vulnerabilities for fewer tokens, the work gets cheaper, it gets faster, and it gets easier to repeat. For defenders, that's really good news. You can scan more code, you can test more systems, you can validate more patches, and you
[13:46] can give smaller teams capabilities that they never had before. For attackers, the same economics are a very dangerous thing. As the cost of finding subtle bugs drops, the number of people who can attempt serious exploitation goes up exponentially. And that's why this story is bigger even than Project Glass Wing or Daybreak as as product launches go. Anthropic's Glass Wing is one response to the cybersecurity threat posed by these advanced models. You give Mythos to trusted defenders and critical software partners so they can find and fix
[14:17] vulnerabilities before less careful actors get there first. OpenAI's Daybreak is another. You push GPT-5.5 and 5.5 Cyber into defensive workflows through trusted access, code ex-security, patch validation, threat modeling, vulnerability triage, and security partners. Different kinds of rollouts, same underlying problem. Models are getting better at finding bugs, and the software world isn't built to patch at that speed. So, in that world where you have Anthropic and OpenAI both rolling out their responses, I think XBEOW's evaluation adds some
[14:47] useful nuance. They found Mythos was very strong at source code audits, native code vulnerability discovery, and reverse engineering, but they said its judgment was still mixed. It can be too literal, it can overstate relevance, it needs validation infrastructure around it to be successful. And I share that because the takeaway from learning more about Mythos and its capabilities should not be fire the security team and let the model do it. It's that the security team is about to need model-assisted workflows to play defense well, and the bottleneck is beginning to shift.
[15:17] Finding bugs is getting really cheap, but validating exploitability, prioritizing fixes, coordinating disclosure, reviewing patches, and deciding which systems are critical is going to take more effort and probably going to take a custom harness designed for cybersecurity. If you own software, especially anything that touches off or payments or browsers or operating systems, cloud infrastructure, anything with healthcare or finance, industrial systems, developer tooling, you should assume that the vulnerability
[15:48] discovery curve has really shifted. Mythos was not a myth. Pick your highest value code bases and run AI-assisted security review wherever it's allowed. You want to be tracking findings separately from your own static analysis so you can see how AI is helping you. You want to require reproduction before anybody panics, right? You can get a patch and disclosure process ready and make sure that somebody owns the question, what happens when the model finds more severe bugs than the team could fix? That is a real question and it may be all hands on deck for your
[16:18] team once you get access to some advanced models. Prepare for that. That is the larger cyber story. What we learned this week is that Anthropic really did make an extraordinary model in Mythos. They are not making it up when they say it would be risky to release it as is. And as security researchers, as security teams, we need to take this chance that we have to prepare for that. We need to prepare for a world where Mythos-like models will be out and loose by December and we should be hardening up our defenses in the
[16:48] meantime. And that's one of the things actually that we need to be using 5.5, which is generally available, and Mythos, when it comes out, to do. We need to be using them really aggressively to scan our systems, check for bugs, and harden up our systems in preparation for the more generally available open weights models that we know will be coming at Mythos-level capability in about 6 months or so. Our fifth and final story is AWS WorkSpaces for AI agents. This one might sound boring, but it actually is one of the
[17:18] more significant stories out there this week. AWS announced that AI agents can now operate desktop applications inside managed Amazon WorkSpaces environments. So, this means a huge amount of important work that still lives inside desktop apps in the enterprise like internal admin consoles, ERP systems, mainframe interfaces, proprietary tools, legacy software, virtualized environments. All of that is now available to an agent in a managed desktop environment. The agent can drive
[17:49] applications the way a human does, but inside an environment with centralized permissions, logging, auditing, screenshots, and metrics. And that changes the pathway to automation for a lot of enterprises. A company doesn't have to wait for every old system to become agent native. The agent can just drive the software that already exists. That's especially relevant in regulated industries and back-office operations. Stuff like claims processing, trade settlement, candidate screening, internal finance ops, healthcare admin, insurance workflows, procurement, legacy customer databases. The places where the
[18:20] work is valuable, but it's also super repetitive and it's often trapped behind old interfaces that no one bothered to update. But, there is a real warning here. Desktop automation is powerful precisely because it bypasses clean integration boundaries. And that's also why it's risky. If an agent clicks through a desktop app, you need to know what it did, why it did it, who authorized it, what data it saw, how to stop it. A screenshot log is useful, but it's not the governance model, right? And so, the practical advice when you
[18:51] look at a capability like this is simple. Don't start with right access. Start with read only or start in draft mode. Let the agent collect information. Let it prepare a form. Let it draft a recommendation. Let it reconcile records or flag exceptions. But, put a human at the final commit point. Once the workflow is stable, then you can let the agent take actions directly. And the story here is not that every legacy app is suddenly solved, it's that the we have no API excuse is getting weaker and weaker and weaker by the month, and agents are going to reach
[19:22] more of the old enterprise software stack faster than people expected. So, backing out, what are you going to do with these five stories? If your team uses Notion, you have something to do tomorrow, right? You can pick out a database, you can design an agent workflow, and you can get right to work expanding the Notion work surface. If you are using the Claude Code paid plan, the 20-bucker, the 200-buck version, and you have an open claw, this is one of those moments you have to evaluate. Do you want to try to go back to using your claw with Anthropic with Claude Code? A
[19:53] lot of developers liked that best when open claw first came out, or do you want to stick with whatever you've moved to since? In many cases, it's OpenAI's ChatGPT 5.5. I don't have a clear answer for you. My my sense from talking to developers who tend to be cost-sensitive is that they will go with the plan where they don't have to do math in their head, and the plan that is simplest for that right now is OpenAI's approach to the open claw and side-party agents using your plan. It's just it's easier to say, "Yep, you can do it, and it's
[20:24] not going to be a problem," versus Anthropic making you do math on how much you can do and what your cap is and when the API billing kicks in. Simple math wins. And if you're used to using Claude to drive your open claw or to use agents, you have to do some math this week, right? And most people don't like to do math. Even most developers don't like to do math in my experience. You have to decide if you think that the credits that you're going to get under Anthropic are going to be enough to move you to keep your open claw or agent on Claude Code, or if you think it's worth
[20:56] it to move somewhere else, or if you think you'd prefer to stay somewhere else if you moved already when Anthropic cut off third-party access entirely in April. Now, Now you're an open claw user or an agent user, and you previously powered your Open Claw with Claude, you have some math to do this week. You have to decide if Anthropic's shift in allowing some agent use back into the plan is enough to bring you back, or if you want to stick with whatever you went to. In many cases, folks have stuck with the new Open AI plan, which is a very
[21:27] simple, transparent way to handle billing for Claw agents. And I don't know. I don't have a clear answer for you. It's going to depend on the model you're using. It's going to depend on the task your Claw does. It's going to depend on your understanding of how token heavy those assignments are, and how often you're going to run into limits. And so there's not a one-size-fits-all solution. Instead, you need to audit and look at how your Claw, your agent is handling token usage, and do some calculations from there. And also, hopefully, benchmark the performance of
[21:58] your Claw on different models to see how it does. And then you can make a decision. And that's how you, in a responsible way, handle news like what Anthropic dropped. It's not as simple as good or bad. It's a lot about how you use the model, and what models support your workflows and your tasks. And of course, if you're in the software security space, the token efficiency and the prowess of Mythos and ChatGPT 5.5 really need to be a wake-up signal for you. I know that Mythos isn't out to everybody yet, but 5.5 is very good, and it is generally available. So, don't
[22:28] wait. Use AI to help you find vulnerabilities today, and figure out how to prioritize and patch those aggressively. Because the world is changing when it comes to vulnerabilities and software detection, and it's changing really fast. And last but not least, if your company has important work trapped in legacy desktop software, and almost every company does, look at AWS WorkSpaces. Not because every agent needs a desktop, but because some of the most valuable workflows in the company have been stuck precisely because the systems are old and visual
[22:59] and not API first. One of the interesting tipping point moments we've reached is around the ability of AI to use those interfaces. Do you remember in 2025, AI was terrible at this. We laughed at it. We said it was bad and it got good. And that is because of the scaling laws. And so scaling laws have enabled us to get good at computer use to a point where it becomes possible for AWS to launch a product like managed desktops and it's just not a big thing. Think about that the next time you hit a
[23:29] limit for what AI can do because I guarantee you we are going to blow past that limit sooner than you expect. One of the consequences of scaling laws is that we don't understand the impact of exponential growth very well until it jumps up and hits us in the face. And the ability to use desktops, the ability to use software, the ability to jump into the current enterprise stack and say we don't need an API, we don't need you to make an MCP, the agent can just use the software. I think a lot of people aren't fully ready for the
[23:59] implication of that for the rest of the computing world. If AI got that good at using software just like a human does in the last few months, what's next? And that's what I'll leave you with. If you enjoyed this, I'll be back with more news very shortly. Uh subscribe and let's have fun. Cheers.
Resumen de investigación





Resumen — 5 historias de IA de la semana


5 historias de IA que cambian decisiones esta semana

TL;DR

  • Notion lanzó el 13 de mayo una plataforma para desarrolladores que convierte el workspace en programable: workers, webhooks, sync desde Salesforce/Stripe/GitHub/Zendesk/Postgres y external agents API (Claude, Codex).
  • Anthropic volvió a apretar los límites de uso de Claude tras el runaway agentic ramp; OpenAI contraatacó ofreciendo 2 months of free Codex a nuevos clientes business (Axios / respuesta de Sam Altman).
  • Mythos preview de Anthropic lidera la cadena de ataque completa (recon → network takeover) con mejor token budget que GPT-5.5 según el UK AI Security Institute y XBEOW; ambos modelos cambian la curva de descubrimiento de vulnerabilidades.

Las cinco historias

◆ Notion: una puerta de entrada para los agentes al workspace

El 13 de mayo, Notion lanzó una developer platform que va más allá de añadir un botón de IA. Incluye una command line interface, workers hosted en la propia infraestructura (p. ej. database sync desde Salesforce, Stripe, GitHub, Zendesk, Postgres), webhooks, custom agent tools y una external agents API para traer a Claude, Codex u otros al workspace interno como participantes.

El propio locutor lo enmarca en una frase: el headline no es "Notion launched AI", sino que "Notion is expanding on their AI launch to become the workbench where humans and agents share context". El ejemplo concreto: un deal cierra en Salesforce → webhook dispara un worker de Notion → el worker monta el onboarding workspace, trae plan data, account notes, success criteria, milestones y support history → un agente redacta el kickoff plan → el CSM revisa en Notion antes de tocar al cliente. Todo el flujo descrito por el locutor como "foundational components that this release was built for that form a coherent workflow".

▶ Claude se queda sin ordenadores: la trampa del éxito

Anthropic, que "kicked off the agent revolution back in December", tiene el problema inverso: el éxito del agentic ramp. Axios reportó que Anthropic está moviendo parte del uso de agent tools por terceros detrás de su propio credit meter, y Sam Altman respondió ofreciendo "2 months of free Codex" a nuevos clientes business.

La cronología concreta del locutor: en "early April, mid-April 2026", Anthropic empezó a apretar las runaway agentic flows; primero cerró el uso de suscripciones personales ($20 o $200 al mes) por parte de terceros —lo que "killed a lot of projects" y empujó devs a OpenAI, que fue especialmente amable con OpenClaw. Ahora, como gesto hacia la comunidad OpenClaw, permite algo de third-party usage pero con "a rate limit on it... use it or lose it... after that, you have to pay the buy-the-token API billing cost". El mensaje para builders: "the limit... is not just an accounting detail or a billing detail. It's actually part of the user experience".

◆ Anthropic vs OpenAI: cuello a cuello y un dato que cambia la foto

El locutor rompe la narrativa de OpenAI como "the gorilla in the room": "now it is fair to say that Anthropic and OpenAI are neck and neck by most revenue terms". La base que da: ambas cerca o ligeramente por encima de $30B en annualized revenue.

El hard fact que no viene de ninguna de las dos compañías sale de Ramp (que ve cómo las empresas gastan porque procesa sus tarjetas): "for the first time Anthropic has more verified business customers than OpenAI". La razón última que da el locutor, citando a Dario Amodei: "I underplanned for growth, and I'm trying to find the compute to support the sort of growth we've had"; Anthropic "had planned for 10x growth in a year, and they're over 80x". Revenue aquí no es lagging indicator sino leading indicator de la presión sobre compute y la supply chain.

▶ Mythos y GPT-5.5: la curva de vulnerabilidades se desplaza

Dos evaluaciones independientes cayeron esta semana sobre Mythos preview: la de XBEOW y la del UK AI Security Institute. Ambas series de tests cubren la cadena completa: "reconnaissance, credential theft, lateral movement through the attack surface, web app exploitation, privilege escalation, command and control persistence, infrastructure compromise, and finally full network takeover".

Dato clave del AI Security Institute: "Mythos preview gets farther in that attack chain on the same token budget than any other model", comparado con GPT-5.5, 5.5 Cyber, otros Claude Opus, Codex y modelos antiguos. Esto pesa porque GPT-5.5 ya era muy fuerte: el propio OpenAI lo posiciona como más token-efficient que 5.4 y "in OpenAI's own evals, 5.5 is ahead of Opus 4.7 on the cybersecurity benchmark Cyber Gym". El AI Security Institute añade que "5.5 itself significantly exceeds the old cyber progress trend". XBEOW matiza que Mythos es fuerte en "source code audits, native code vulnerability discovery, and reverse engineering" pero "its judgment was still mixed. It can be too literal, it can overstate relevance, it needs validation infrastructure around it to be successful".

Las respuestas de producto ya existen: Glass Wing (Anthropic) y Daybreak (OpenAI) canalizan los modelos hacia defensores y software partners. El takeaway práctico del locutor: "the vulnerability discovery curve has really shifted"; no se trata de "fire the security team and let the model do it", sino de workflows model-assisted y de "prepare for a world where Mythos-like models will be out and loose by December", con "open weights models that we know will be coming at Mythos-level capability in about 6 months or so".

◆ AWS WorkSpaces para agentes: la excusa del "no hay API" se debilita

AWS anunció que agentes de IA pueden ahora operar desktop applications dentro de managed Amazon WorkSpaces environments, con permisos centralizados, logging, auditing, screenshots y métricas. Eso abre al agente caminos donde antes solo había UI humana: "internal admin consoles, ERP systems, mainframe interfaces, proprietary tools, legacy software, virtualized environments" — claims processing, trade settlement, candidate screening, internal finance ops, healthcare admin, insurance workflows, procurement, legacy customer databases.

La advertencia operativa del propio locutor: "Don't start with write access. Start with read only or start in draft mode... put a human at the final commit point. Once the workflow is stable, then you can let the agent take actions directly". La tesis de fondo: "the we have no API excuse is getting weaker and weaker and weaker by the month", y lo achaca explícitamente a las scaling laws — "scaling laws have enabled us to get good at computer use to a point where it becomes possible for AWS to launch a product like managed desktops and it's just not a big thing".

◆ Buscar el alpha

El alpha del episodio no está en trades concretos sino en hacia dónde se inclina la industria de la IA según el invitado: del lanzamiento de modelos al despliegue en artefactos reales (workspaces, desktops, security chains), con compute como cuello de botella y las scaling laws como motor subyacente. Las cifras y nombres citados literalmente:

  • Launched / shipped: Notion developer platform (May 13) con workers, webhooks, external agents API para Claude y Codex; AWS WorkSpaces para agentes, que opera desktop apps en entornos gestionados con logging, auditoría y métricas; Mythos preview evaluado por XBEOW y por el UK AI Security Institute sobre la cadena de ataque completa (recon → network takeover); GPT-5.5 y GPT-5.5 Cyber como base comparable, con OpenAI describiendo 5.5 como más token-efficient que 5.4 y领先 a Opus 4.7 en Cyber Gym según evals propios.
  • Players y roadmap: Anthropic — ejecutando el "agent revolution" que arrancó en diciembre, ahora con compute insuficiente, planeó 10x y va por 80x; Glass Wing como canal hacia trusted defenders. OpenAI — contraataque comercial con "2 months of free Codex" y Daybreak como canal defensivo. AWS — managed desktops como vía para que el agente entre en legacy stack.
  • Capability delta: Mythos preview cubre la cadena de ataque completa (reconnaissance, credential theft, lateral movement, web app exploitation, privilege escalation, command and control persistence, infrastructure compromise, network takeover) llegando más lejos que cualquier otro modelo en el mismo token budget; GPT-5.5 a su vez excede la old cyber progress trend.
  • Adoption signal: Ramp reporta que "for the first time Anthropic has more verified business customers than OpenAI"; ambas cerca o ligeramente por encima de $30B en annualized revenue.
  • Catalysts / inflections que el locutor vigila: cuándo los open weights models alcanzan capability tipo Mythos (~6 meses); la release general de Mythos y "Mythos-like models... out and loose by December"; la decisión de los developers OpenClaw sobre si volver a Claude Code con la nueva rate limit mensual o quedarse con el modelo OpenAI.
  • Predicciones con horizonte: "Mythos-like models will be out and loose by December"; "open weights models... at Mythos-level capability in about 6 months or so".
  • Calls contra-consenso: Anthropic está neck and neck con OpenAI por revenue y ya le supera en verified business customers según Ramp; las scaling laws ya permiten computer use serio (AWS WorkSpaces es prueba), por lo que "the we have no API excuse is getting weaker and weaker and weaker by the month"; la vulnerabilidad discovery curve ya se ha desplazado y encontrar bugs se vuelve barato, lo que cambia el bottleneck a validación, disclosure y patch.
Activo / señal / lectura
Activo Señal Lectura
BTC (Bitcoin) Un usuario recuperó 5 BTC (~$400,000) de un wallet.dat de hace 11 años gracias a Claude, no a un cracker. El alpha implícito del propio locutor: "the model launches are still happening, but the more interesting stuff is quieter and much more specific. Agents are starting to do real work on real artifacts... and it's beginning to change real life".
AMZN (AWS) AWS WorkSpaces para agentes: desktop apps operables dentro de managed environments con logging, auditoría y permisos centralizados. Las scaling laws vuelven creíble la computer use; "the we have no API excuse is getting weaker and weaker and weaker by the month". Back-office y regulated industries pasan de estar atrapadas en legacy UI a estar expuestas al agente, con el guardrail de read-only/draft + human commit.
CRM (Salesforce) — vía Notion Notion database sync desde Salesforce (más Stripe, GitHub, Zendesk, Postgres) mediante workers hosted. El workspace donde los humanos ya trabajan se vuelve programable: webhook → worker → agente → revisión humana, todo en Notion. Cambia la pathway a agentic ops sin re-platforming.
La vuelta de tuerca: el invitado cuenta cinco historias y en todas el alpha está en lo aburrido: límites de suscripción que se rompen, una developer platform de un workspace, managed desktops y eval de un modelo de cyber. Lo que está describiendo sin enunciarlo así es que la fase de la IA ha pasado de "qué modelo lanza qué capacidad" a "qué compute, qué pricing y qué governance hacen que el agente llegue al artifact real" — y que la compute shortage de Anthropic (10x planeado, 80x real) y la calidad de Mythos en token efficiency son los dos hechos que más pesan sobre todo lo demás, incluido el roadmap de open weights a ~6 meses y la presión competitiva con OpenAI en revenue y en business customers.


Generado con algoritmo v2.1-anchor-first · modelo MiniMax-M3 · 2026-07-05T19:34:53Z

← Volver al listado de vídeos

Scroll al inicio