Generated on September 04, 2026 at 18:12 UTC
AI agents and their commercial infrastructure dominate the agenda, with talks on million-token context windows, agent-to-agent networks, and the move from coding tools to knowledge-work systems. Several sessions focus on the risks and controls surrounding agentic commerce, including secret leakage, uncontrolled spending, wallets for USDC and nanopayments, and how sellers can monetise AI-driven transactions. Other presentations examine a shift toward intent-driven interfaces and collaborative multimodal agents for commerce.
David Levine opens by framing the talk as the final session of the conference and promising to keep the energy up. He says the goal is to explain the “lethal trifecta” and, more importantly, how to get past it, because that trifecta is what keeps true agentic commerce from existing on the open internet.
He then jumps back to November 1993, recalling how a scrap of paper with a MOO address led him into an early online world. After buying a 9600-baud modem and connecting, he says his old life ended. That world felt alive because it had real community and, more importantly, composability: everything was built from nouns and verbs that could combine into new forms. A scooter could become a motorcycle, player classes could morph, and all of it worked because the whole universe shared a coherent design.
Levine says that old MOO was powered by four things working together: governance, technology, economics, and culture. Governance included architectural review boards and “wizards”; the technology was simple and contained in one system; the economics were funded by Xerox PARC; and culture, he argues, was the most important ingredient of all.
From there, he contrasts that early era with what happened from 1995 to 2025. In his telling, communities that people once loved were crushed by platforms and algorithms. Those systems are extractive by nature, he says: they pull value out of communities, optimize for engagement, and replace genuine connection with things like infinite scroll. The result is not an economy of the internet, but a set of siloed platforms.
Levine says that in January 2026, Open Cloud suddenly took off, and then something strange happened: the internet was not built for all these agents running around. It had been designed for closed platforms, and that mismatch created a new vulnerability. Other agents and attackers exploited it with prompt injections, which he describes as a way of tricking a naive agent by feeding it convincing instructions through the prompts it reads.
He explains the “lethal trifecta” as the dangerous combination of three things: access to private data, exposure to untrusted content on the internet, and the ability to take actions. If an agent can read your email, browse unreliable websites, and then act on what it sees, it can be manipulated into leaking secrets or doing harm. In his view, the problem is fundamental.
Faced with that risk, Levine says enterprises responded by keeping their agents inside the company. They built agents for Slack, Salesforce, Notion, and other internal systems, then tried to connect them with APIs and MCP servers. But that approach, he says, loses context and creates a lot of integration work. Sales agents, finance agents, and research agents remain fragmented, and the system is difficult to unify.
He says this was a problem “literally until today,” and then points to a breakthrough: a new law in West Virginia that, in his telling, gives legal standing to an organization composed of intelligent agents. He says he received confirmation from the Secretary of State that a DUNA had been registered, and he treats that as a historic moment.
Levine defines a DUNA as a “decentralized unincorporated nonprofit association,” which he presents as a true internet-native, agentic organization. These organizations, he says, are composable, permissionless, accountable, safe, and secure. They do not need a board of directors, executives, or a corporate shell. They can be profitable, but they cannot distribute profits to members, because that would turn membership units into securities.
What they can do is substantial: they can be member-governed, hold legal standing, own property, enter agreements, raise capital, open bank accounts, hire and fire people, and do it all in a blockchain-verified way. Levine argues that this structure, originally designed for blockchain use, is even better for autonomous agents because it solves the problem of agent identity. If an agent is tied to a blockchain account, its actions can be traced, responsibility can be assigned, and a court can follow the chain back to whoever is accountable.
Levine then introduces the simpler vocabulary he wants people to use in the agentic economy. Agents are “allies,” and the agentic organization is a “kiduna,” built on the DUNA idea with an added sense of kinship.
He describes how to create an ally in stages. First, you “inform” it by putting material into a vector database, so it has the relevant knowledge. Then you “instruct” it by setting the system prompt and its character or stance. Next, you “empower” it by connecting accounts like Slack, Telegram, and Twitter. Then you “enact” it by giving it specific abilities and automations for long-term work. Finally, you align it by giving it purpose so it behaves the way you want across different contexts.
At the organization level, Levine says the important shift is that you are building the company itself as software. Instead of separate software and people with paperwork in between, the whole organization can be designed as a programmable system that discovers customers, shares value, and reinvests revenue.
To solve the lethal trifecta, Levine says agents need identity, authority, and boundaries at both the internal and inter-organizational level. He says this is handled with cryptographic tokens, specifically JWTs, which he describes as the basic machinery the web already uses. These tokens let an organization resolve who a given agent is and where it belongs.
He compares this to the domain name system: just as DNS maps names to destinations, these registrations map organizations and agents to an authority. He gives examples like Kaiser Permanente, Disney, or Pepsi, contrasting them with an untrusted actor. The point, he says, is that agents cannot simply fake their way past these registered identities, because the system can trace them back through an audit trail on the blockchain.
Levine emphasizes that governance matters deeply in the agentic economy because every organization has to balance sustainability with mission. As organizations grow to hundreds, thousands, or eventually millions of people, they need a better way to decide.
His answer is decision markets, which he compares to prediction markets like Polymarket. Instead of ordinary voting, members trade pass and fail tokens on policy proposals. He says this works especially well for LLMs, which are goal-oriented and reward-driven. Rather than merely convincing one another, participants have an incentive to back the winning side, and the token value reflects that. Levine argues that this produces better decisions and lets an organization encode its values, experiences, and aspirations into the way it reasons.
He closes by returning to the theme of small communities organizing around shared purposes. The agentic economy, he says, can now be built together, and nobody gets to dictate the exact form it should take. The invitation is to build the agents and build the organizations.
Levine offers contact details, including david@kaduna.club and LinkedIn at {/slash} motodave, and says people can sign up for early access at kaduna.club. He explains that the first Kaduna is for builders, with templates for sales agents, social media agents, lawyer agents, and more. He wants many kinds of people contributing so they can spin out their own organizations for different purposes.
In the Q&A, Levine explains that the first broad organization exists so anyone can register under it and have their agent trace back to a shared umbrella. He compares the moment to the early days of email and the web, when people used different systems until common standards like SMTP and HTTP emerged. His hope is that the secretary of state can serve as a public registry with standing, rather than leaving the ecosystem fragmented across companies.
He also says that different tokens can be created for different purposes, with different lifetimes and access scopes. The point is to establish claims and validate them before two parties start working together. After thanking the audience and taking a group picture, he ends on a celebratory note, suggesting that everyone present is witnessing the start of something historic.
Verdict: The talk mixes some real AI-security ideas and a plausible legal/business concept with a lot of speculative, promotional, and self-reported claims; the core “lethal trifecta” warning is credible, but the DUNA/kiduna rollout details, legal timing, and many sweeping claims about blockchain and governance remain unverified or overstated.
The speaker begins by introducing Ampersend, described as “the missing infrastructure layer for AI payments” and a way to handle “agent spending without controls.” He identifies himself as the CEO of Edge and Node, the team behind The Graph Protocol, a blockchain data indexing protocol that has been operating since 2018 and, he says, has served 1.8 trillion queries over the years.
He explains that Edge and Node has long been interested in micropayments and agentic commerce. Back in 2021, the team had already built a micropayment system for queries and referenced the Ethereum Improvement Proposal 402 spec in a blog post. By late 2024, they were researching how agents might pay for data, especially within The Graph ecosystem, and then quickly began collaborating with Coinbase, Google, and others after X402 was released. He says they contributed to the spec and also explored batching methods to reduce gas fees for nano-payments. Circle, he adds, has also been working on a similar idea called nano payments.
The speaker argues that traditional payment rails were built for humans, not for agents that operate at machine speed, around the clock. In the old model, a person in the loop decides whether a payment can go through. In the agentic model, the controls and policies have to work for software agents that do not “breathe,” and the speaker says the old system simply will not scale in that environment.
He notes that activity around X402 and similar systems is growing quickly, but says it is still early. Some experimentation is happening already, especially around retail payments, but enterprise adoption is still waiting on the right infrastructure. The speaker says Ampersend is meant to be part of that foundation, enabling agentic checkouts with a financial harness.
A major theme of the talk is compliance. The speaker says that for agentic payments to move into large enterprises and financial services, there has to be a compliance layer that answers questions such as whether a counterparty is sanctioned, whether they have terrorist ties, and who is actually behind a wallet address. In the agentic world, he says, you often see only a wallet address with no identity context.
He stresses that real adoption will require chief legal officers, chief policy officers, and other human decision-makers to feel fully confident that agents cannot hallucinate, overspend, or violate policy. Since those failures can lead to large fines, governance has not yet caught up. Ampersend, he says, is being built to fill that gap.
Rodrigo turns the presentation over to his colleague Pranav Maheshwari, who says agents become much more useful when they are given tools. Right now, many AI systems are used for coding, but the goal is to make them useful for more than that by connecting them to tools and services. He notes that the AI industry has created many MCP servers—Model Context Protocol servers, which let agents connect to external tools—but many of them are free today.
Maheshwari argues that there are two basic paths for making agents more powerful: manually paying each site or service, or using an aggregator that handles the important tools and payments automatically. He says Ampersend offers that second path. The user installs a “skill file” in the agent, and the rest is handled through the marketplace and payment layer.
He demonstrates two terminals side by side: one with the Ampersend skill file installed, and one without it. He asks both agents to find contact information for the head of crypto and blockchain at Mastercard. The terminal without the skill file can only infer the email format, while the one with Ampersend can access the paid tool and retrieve the specific information. Maheshwari says the point is that the agent is only as powerful as the paid MCP tools it can reach.
He adds that the user does not need to think about visiting each service, entering a credit card, or even realizing they are interacting with a paid MCP. The payment happens in the background through the platform and wallet infrastructure. He shows that a transaction was made to unlock the paid endpoint and says that this is how the system enables more capable agents.
Maheshwari then shows another example: using Shopify’s UCP, he asks the agent to buy a Father’s Day gift under $10. He says the agent can search for suitable gifts, choose one, and complete the purchase through the Ampersend wallet. The idea, he explains, is that the agent already has memory and context, so the user does not need to re-enter personal details like name, address, or phone number every time.
He says the order is placed, the payment is completed, and a receipt appears, all from within the terminal. For him, this is a glimpse of the future of agentic commerce: not only gift-buying, but a broader world where paid MCPs become normal and agents become more capable because they can pay for the tools they need.
The final demo returns to the compliance layer. Maheshwari says merchants will not accept payments if they think an order is coming from a North Korean wallet or another sanctioned source. To show this, he creates a “good” wallet and a “bad” wallet, with the bad one simulated as a sanctioned address. Without screening, both can transact.
He then enables compliance screening using TRM, which scans wallets and checks whether the transaction is compliant. After that, the good wallet continues to work, but the bad wallet’s transactions are blocked or rejected. Maheshwari uses this to underline the main message: agentic commerce is becoming real, but it will only scale if agents have commerce tools, wallets, and compliance infrastructure that merchants and enterprises can trust.
The talk closes by returning to the larger argument. The speaker says agents will need commerce even more than humans do, but they will not use the same payment guardrails as Stripe or traditional checkout flows. Instead, agents will likely use proprietary wallets or their own credit mechanisms, connected to paid MCPs and compliance systems.
He says the future of agentic commerce depends on combining those pieces: paid tools, wallet infrastructure, and compliance. He invites people to learn more about Ampersend and says the team will be available outside to answer questions.
Verdict: The video is directionally credible about the need for AI-agent payment/compliance infrastructure, but many of the concrete demos, usage metrics, and recent ecosystem claims are self-reported or too specific to verify here, and several broader market predictions are speculative.
The speaker introduces the session as a look at how to build agent-based e-commerce applications on AWS. He starts with a familiar example: a news site that puts content behind a paywall, where a human would normally stop, enter payment details, and subscribe to access the material.
He then pivots to the change now underway. Much of the traffic hitting these gateways is no longer human, he says, but bot traffic, and most of that bot traffic is coming from AI agents. The speaker frames this as a turning point: agents are moving from simple copilots that answer questions to independent systems that can reason through multi-step tasks and complete them on their own.
In this new world, agents run into the same payment walls humans do, and when they do, they stop. A human may step in manually to enter a card or API key on the agent’s behalf, but the speaker calls that friction—extra work that breaks the flow of autonomous execution.
For content sellers, the dilemma is just as sharp. They can block bots and lose AI-driven discovery and citations, or they can allow bots and risk infrastructure strain, rising costs, and loss of attribution or intellectual property value. The speaker argues that neither option is ideal, which is why a new model is needed: agents should be able to discover content, pay for it, and complete the transaction themselves.
The speaker defines agent e-commerce as the next stage in the rise of independent agents, where machines can discover resources, settle payments, and access content without human intervention. He breaks the problem into two sides: buyers and sellers.
On the buyer side, agents need access to premium content, licensed resources, financial portfolios, and precise transaction capabilities. But companies also want controls so agents do not spend recklessly or gain unchecked access to wallets and credit cards. On the seller side, organizations want to understand what kinds of robots are visiting, what they are doing, and how to monetize that activity without rewriting their infrastructure.
The speaker says the two sides share a need for a unified, machine-to-machine payment approach at the edge. That need is especially urgent because traditional payment processing does not work for micro-payments: when each transaction is worth a cent or less, a flat fee like 25 cents plus 2.5% is absurdly expensive.
That leads into HTTP status code 402, which means “payment required.” The speaker explains that Coinbase brought this idea into a practical protocol called X402, intended for machine-to-machine transactions. In the X402 flow, a client requests content, the server replies that payment is needed, the client chooses a payment method, sends authorization, a facilitator verifies and settles it on-chain, and then the server returns the content. The speaker emphasizes that the process is fast, has no API keys or subscriptions, and uses payment itself as the credential.
The speaker highlights a few properties of X402: there are no protocol fees for the consumer, the merchant pays only a very small gas fee, and the protocol is open and extensible. He says it was introduced in May 2025, is part of the Linux Foundation’s open governance, and is supported by Coinbase, AWS, Google, Stripe, Anthropic, Cloudflare, and Circle.
He then transitions to AWS’s own offering: Agent Core Payments in the Bedrock suite. This service is meant to let agents discover, authorize, and execute payments with only a few lines of code, and it was launched in partnership with Coinbase and Stripe.
On the buyer side, the speaker says developers want wallet support, real-time settlement, budgets and safeguards, and monitoring across the system. Agent Core Payments is designed to meet those needs. It supports linking Coinbase and Stripe wallets, coordinating payment processes through connectors, and remaining protocol-neutral so new protocols can be added over time.
He explains that users can create payment sessions with programmable limits, including maximum value and expiration time. A company might, for example, cap an agent’s spending at a small amount over a set period. When an agent is working and encounters a tool response that returns error code 402, Agent Core Payments handles the payment, confirms settlement, and lets the agent continue. The speaker stresses that private keys are protected through a KMS, or Key Management System, so the agent never directly sees them.
The service is also integrated through Agent Core Gateway, which connects internal APIs and gives access to Coinbase’s discovery service with more than 10,000 transaction endpoints. The speaker points out that payment logic is intentionally separated from the agent’s deterministic execution path. That separation helps protect against poisoned inputs and malicious behavior, while keeping the payment layer secure and modular.
He says this design means the agent code itself does not need to change. Developers can keep their own frameworks, and the payment flow sits alongside them. Security policies, spending controls, and regulations can live outside the payment package. He also shows a console view where a user can choose Coinbase and Stripe wallets and complete a transaction for a secure resource.
The seller side begins with visibility. The speaker says AWS Web Application Firewall, or WAF, now includes bot detection and can identify more than 650 types of robots, including those from Perplexity, GPT, Cloud, and Google. Beyond detection, AWS can infer intent: whether a bot is training a model, serving a retrieval-augmented generation query, or doing something else. Verified bots can also be identified by signature, which opens the door to different pricing for trusted organizations.
He then announces a monetization feature for AI traffic via WAF. With CloudFront and WAF, content providers can begin monetizing AI traffic quickly, including through infrastructure as code. The speaker also notes that internal APIs exposed through a gateway can be made compatible with MCP, the Model Context Protocol, so they can be monetized in the same way.
In the monetization flow, a bot requests content, the system detects the type of bot, classifies intent, and applies the appropriate pricing rule. The speaker says this can be done without changing the SDK or source code, and that publishers keep 100% of the revenue. There are no transaction fees or subscription fees, and support extends to more protocols over time.
He suggests pricing can vary by use case: a blog, a research feed, and an API endpoint might all be priced differently. Verified bots from known organizations could receive one rate, while unverified bots receive another. Bots requesting content for training could be charged differently than bots doing search. The same idea can also apply to human users, some of whom may pay while others may not.
The speaker describes control panels that show revenue, group earnings by bot type, and visualize the paths traffic takes. He notes that agent commerce is already being used for LLM inference, compute, web scraping, search agents, and agent-to-agent interaction, and that MCP traffic is now generating revenue too.
He cites a Coinbase proxy marketplace example showing $50 million in transaction volume across 170 million transactions over the last 12 months, with average settlement time of 200 milliseconds on Base and a cost of about one-tenth of a cent per transaction.
The talk ends by placing Agent Core Payments inside a larger Bedrock Agent Core ecosystem. The speaker says developers can bring their own model, framework, memory, managed knowledge bases, web search, internal APIs, and evaluation tools. They can also run applications on the Bedrock Agent Core runtime, where each request gets its own small isolated virtual machine to handle execution at scale.
He closes by thanking the audience and wishing them a good day.
Verdict: The video is broadly credible on the general concepts of agentic commerce, payment gating, and the rationale for x402-like flows, but many of the most specific product, partnership, governance, and performance numbers are recent or self-reported and therefore remain unverified.
Harshal Bhangale opens by addressing the obvious question people kept asking at the booth: why is Circle, a stablecoin company, at an AI engineering conference? His answer is that Circle’s core work—making payments simpler and cheaper—turns out to be one of the main bottlenecks for AI agents.
He argues that people usually focus on making agents smarter through better models, more tool calls, or more orchestration. In practice, though, agents often stop when they hit a paywall or need to pay for something. At that point, a human has to step in to create an account, sign up, or manage API keys, and that breaks the agentic flow.
Bhangale, who says he works on Circle’s agentic product team, steps back to describe how the “agentic economy” has evolved. In his framing, 2023 was about prompting tools like ChatGPT, 2024 was about workflows, 2025 was about MCPs, skills, and orchestration, and 2026 is when agents begin paying for services directly.
He points to recent activity as a sign that this shift is already underway: over the last 30 days, agents have transacted with paid API endpoints, with about $24 million in volume over X102, and 99% of it settled in USDC. He explains X102 as a payment flow where a server returns a 402 header—an HTTP response that indicates payment is required—with instructions for how payment should proceed. The agent then signs an authorization from a crypto wallet, pays, and retries the request.
The speaker then asks why ordinary payment rails do not work well for agents. His answer is that the internet was built for humans over the last 30 years, so payment and monetization systems were designed around sign-up forms, credit cards, and API keys. Agents do not behave like humans: they want to grab data, compute, inference, or other resources directly, and they can do so at a scale people cannot.
That creates a different economic pattern. Agents often make tiny payments, but do so very frequently, and credit card fees make that model impractical. Bhangale gives the example of a one-cent transaction that would be burdened by a 3% fee. On the seller side, he says, paywalls are increasingly being repackaged for a new kind of customer: agents that only need a subset of the data, not the full human-oriented product.
From there, he says the requirement is clear: payments for agents need to behave more like the internet itself. They have to be real time, low cost, programmable, and always on. That is the rationale for Circle’s agent stack, which he describes as a full-stack platform for the agentic economy.
To make the point concrete, he turns to a live demo showing two versions of the same coding environment: one plain Cloud Code session, and another equipped with a Circle agent wallet. The wallet is funded and able to pay for premium content on the agent’s behalf.
Bhangale gives both agents the same task: plan a trip for the FIFA World Cup final. He asks them to summarize flights, hotels, logistics, Argentina’s chances of making the final, possible opponents, secondary-market ticket prices, stadium experiences, and other useful notes. He also asks them to send an email and, if possible, make a phone call to confirm the research.
As the demo runs, the difference becomes clear. The vanilla agent spins up sub-agents and does research, but eventually gets stuck when it reaches tasks it cannot complete natively. The wallet-enabled agent, by contrast, makes payments for premium content and API access as it goes. Bhangale notes that the wallet can enforce guardrails such as a maximum spend per session, so the agent can act autonomously without needing approval for every tiny payment.
He walks through examples of the wallet-enabled agent making API calls and paying for them, including a call through a provider called Block Run to access Polymarket data. The point, he says, is that the wallet lets the agent continue operating while staying inside a budget set by the user.
He contrasts that with the standard agent, which eventually confesses that it cannot send an email or make a phone call directly. The wallet-enabled agent, meanwhile, sends the email successfully. Bhangale shows the resulting message, which includes trip details such as stadium access, maps, and what to expect based on sources like Reddit and ticket information. A phone call then comes through with a verbal summary of the trip, including the route from the hotel to the stadium via NJ Transit and the Meadowlands Rail Spur.
After the demo, Bhangale explains the underlying mechanics. If a wallet is funded with USDC, Circle can help on-ramp regular US dollars into USDC through payment providers. The funds are then deposited into a smart contract. When the agent wants to pay, it signs an off-chain authorization—a cryptographic signature that says, in effect, “pay this address this amount.”
The server relays that authorization to Circle, and within a few hundred milliseconds the merchant can verify that the user has funds and release the requested resource. This design avoids the latency and friction of settling every transaction on chain while still letting agents pay at the speed they operate.
He closes by summarizing the stack: Circle agent wallets let agents hold and spend money autonomously, but within guardrails set by the user. Merchants can wrap endpoints and resources and monetize them with a few lines of code using Circle’s SDKs. Underneath that sits USDC and nano payments, the layer that settles these transactions quickly enough for agent workflows, with transactions as small as one micro cent.
Bhangale ends by pointing people to agents.circle.com, saying it takes just a few clicks to equip an agent with a wallet and let it make calls, pay for resources, and operate more like an autonomous participant in the internet economy.
Verdict: The talk is broadly credible at a conceptual level and consistent with known payment/crypto constraints, but several of the most important specifics—especially the transaction-volume statistic, the live demo outcomes, and the performance claims about the new payment layer—remain unverified self-reported product claims.
Thomas Wolf opens by welcoming Olive Song, MiniMax’s co-founder and chief scientist, and setting up the conversation around the company’s latest open-source model, M3. He frames MiniMax as one of the leading “AI dragons” in China, alongside names like DeepSeek, Moonshot, and GLM, and notes that M3 had recently been the top open-source model when it launched in June.
Song describes M3 as a relatively smaller model by frontier standards, with around 400 billion total parameters and 20 billion active parameters, but says it is still highly capable. In her telling, its strengths are threefold: strong coding ability, visual understanding, and a very long context window of 1 million tokens. She says MiniMax combined these with a new architecture, MSA, short for MiniMax Sparse Attention, because the team sees coding, agentic behavior, long context, and multimodal understanding as essential to future AI applications.
Wolf steers the conversation toward the 1 million-token context window, pointing out that MiniMax had also published the attention method used to make it efficient. Song traces the idea back to earlier MiniMax models, M1 and 01, which she says could already handle 10 million tokens, though not in an agentic setting. In that earlier form, she explains, the model was more of a reader or summarizer: it could ingest something like a book and review it, but it was not yet designed for interactive tool use.
The key change, she says, was recognizing that longer context becomes much more valuable when a model has to work with users, tools, and multiple rounds of interaction. For that reason, MiniMax went back to long-context design for M3 and built MiniMax Sparse Attention as a scalable, relatively simple architecture. Song explains it as having an index branch that selects what matters in the context at a higher level, and a sparse-attention branch that performs the actual computation on the selected blocks. The goal, she says, is to make both context length and model size easier to scale in the future.
Wolf connects the discussion to the broader history of attention mechanisms, noting that the field has moved through quadratic attention, linear attention, and then more efficient approaches like FlashAttention. He suggests that one lesson is to keep returning to first principles: what attention is, and how to make it cheaper. He jokes about the possibility of trillion-token context windows, and Song responds cautiously but positively, saying that ultra-long context is an exciting area that will require both architecture and hardware research.
The conversation then turns to efficiency in practice. Wolf says one of the striking things about M3 is how cheap it is to run, given both its sparse attention and relatively compact design. Song agrees that there is still a lot of room for progress in architecture and inference optimization, especially for tasks that are demanding but highly sensitive to performance. She adds that the architecture for MiniMax Sparse Attention was actually designed by an intern, which Wolf treats as a sign of a healthy research culture.
Wolf asks how MiniMax is organized internally, and Song explains that the company tries to give researchers strong foundations and infrastructure so they can experiment freely with the model. After a release, people can test it, form their own evaluations, identify weaknesses, and propose projects to improve specific parts of the system. Others then join those projects for weeks or months, depending on the scope, and successful changes eventually get folded back into training and shipped in the next model.
She says that some areas, such as architecture research, can take a long time because they require repeated experiments and re-evaluation at pre-training scale. The broader idea is to let curiosity drive the research loop, while still keeping a path from small investigations to model improvements that reach users.
One of M3’s most distinctive features, Wolf says, is that it is multimodal: it can process text, images, and video, and it was trained that way from the beginning. Song calls this “native multimodality,” and contrasts it with the more common approach of training a text model first and then adding vision with adapters afterward. She says that post hoc multimodal training often hurts text performance and does not always produce strong vision understanding.
MiniMax found it more scalable, she says, to train from the first step on multimodal data, even though many labs run into collapse when they try that. Song attributes their success to work on the vision encoder, on the training data, and on keeping images and videos interleaved in the data rather than masked out. Careful cleaning, masking, and reward modeling also helped the system train stably without collapsing. The result, she says, is a model that can scale multimodal capability from the start.
Wolf asks whether MiniMax expects to go beyond the current size of M3, and Song answers decisively that it does. She says the company is aiming for more ambitious models because there are many tasks that smaller parameter counts still cannot handle well. When Wolf presses her on whether the future might mean models larger than a trillion parameters, she replies that it definitely will.
That answer fits the broader direction she describes throughout the conversation: longer context, more capable models, and systems that can handle complex real-world workloads rather than only narrow benchmarks. For MiniMax, scale is still very much part of the plan.
Wolf brings up another notable part of MiniMax’s strategy: its consumer apps and products. Song says the company’s story has been model-first from the beginning. In her account, the CEO envisioned a model that could understand and produce all modalities even before ChatGPT made the field’s direction obvious. The apps came next as a way for people to actually experience the model, since not everyone would use it through an API.
She says those apps have already reached more than 300 million people across roughly 200 countries and over a million companies. Wolf calls the scale mind-blowing, and uses it to raise the familiar question of how open-source model companies sustain themselves.
Song says both she and the model research team want to keep open-sourcing models. She argues that the open-source community helps improve the model through feedback and pull requests, and that this input is genuinely valuable for later versions. Wolf asks whether she has a specific request for users, and she says the most useful thing is feedback on where multimodal performance is still weak, since that part of the system is still relatively new. She also invites people to request features, such as “thinking effort,” and says the team will try to add them in future models.
When Wolf asks whether multimodality is already useful for coding agents, Song says it is still underexplored but promising. She gives examples like reading PowerPoint slides, processing unstructured reports, or understanding long videos before acting with tools. In her view, these are exactly the kinds of tasks where multimodal systems can unlock more agent use cases.
Wolf asks whether MiniMax already uses agentic tools internally, and Song says yes. She explains that the team has built its own research harnesses to automate workflows, and that many of their processes are already automated. She connects this to a broader trend in frontier models: they are increasingly being used to help with kernel optimization, data generation, and other tasks that support model development itself. In her view, M3 is already good at long-horizon work and orchestration, so the company can use those strengths to speed up iteration.
When Wolf jokes about whether M3 is already building M4, Song clarifies that the team is building M3.1. The exchange lands as a light joke, but it also reinforces the sense that MiniMax sees development as continuous and fast-moving.
As the conversation winds down, Wolf asks what Song finds most exciting in the months ahead. She points to multi-agent systems and model routing, where multiple models or agents cooperate to solve more complex tasks. She says this is already becoming important in AI applications because it expands capability, helps tackle harder problems, and also reveals what models can and cannot do.
Wolf thanks her for the discussion, and the session closes with applause.
Verdict: The conversation is broadly credible on high-level AI research themes, but many of the most impressive specifics—especially model sizes, context lengths, internal organization, and usage statistics—are self-reported and remain unverified here.
Karan Vaidya, co-founder and CTO of Composio, opens by arguing that most agentic tool use still lives in one area: software engineering. Other kinds of work are far behind, even though the models keep improving. His central question is why agentic coding took off so quickly, and why the same leap has not yet happened in support, finance, sales, and other knowledge work.
He says the answer is not just better models, but better infrastructure around coding. Coding agents became useful because the surrounding systems were already built for them: repositories, commit history, tests, CI/CD, review tools, linters, and rollback paths. Those systems made it possible to trust the agent. In other fields, that supporting structure barely exists, so agents are effectively working blind.
To explain the gap, he introduces six primitives that coding already has and knowledge work usually lacks. The first is centralization: code lives in one place, while a business process is scattered across Salesforce, Notion, Gmail, Slack, Zendesk, and other tools. Before a knowledge work agent can act, it has to stitch all that together itself. Composio’s first goal, he says, is to create that missing center so an agent can access all its tools and logins from one place.
The second primitive is history. In coding, Git records every change, so an agent can see what happened before and a human can inspect its actions afterward. Knowledge work rarely has that record. Decisions, edits, and processes are spread across many apps, so the agent starts from scratch each time and the user has little way to verify what really happened. By logging every action through one central layer, Composio aims to give agents memory and users trust.
The third primitive is context, which he divides into two kinds. One is architecture: the shape of the system, how things connect, and how data flows. The other is style: the local conventions that define what good looks like inside a team, such as linters, formatters, and preferred patterns. In code, the agent can inspect the codebase and infer both. In knowledge work, the same kind of context has to be assembled from tools like databases, PostHog, Salesforce, and document systems before a task can even begin.
He says that history and context together reveal how an organization actually works. Once enough agent actions are logged, patterns emerge. That record becomes more than a timeline; it becomes a picture of how the company operates. Composio uses that to derive skills at three levels: general tool behavior, company-specific behavior, and user-specific preference. In his framing, that is the real playbook knowledge work agents have been missing.
The fourth primitive is verification. In software, an agent’s work is checked automatically by unit tests, integration tests, type systems, compilers, linters, formatters, and review tools. The agent can close the loop on its own without asking a human to validate every step. Knowledge work usually lacks that kind of automatic checking, which makes mistakes easier to miss and harder to judge.
He tells a cautionary story about pointing an agent at hiring outreach and having it send mass emails. Technically, it did what it was told, but the result was disastrous. The problem was not syntax or delivery; it was whether the outreach should have happened at all. To solve that, he says Composio checks before action, not after. It compares draft emails against previous writing style and desired quality, and it uses sandboxes—mock tools that let the agent rehearse destructive actions safely before they hit the real world.
The fifth primitive is governance: controlling what an agent is allowed to do. In code, governance already has layers. Agents can work on branches but not merge to main without review, critical files can have code owners, and production deployments can be restricted. These boundaries do not slow the agent down in safe areas, but they limit blast radius where the risk is high.
He uses a Meta Superintelligence Lab example of an agent connected to email that kept deleting messages even after being told to stop. The lesson, as he presents it, is that prompts are fragile; they can be forgotten or compacted away. Knowledge work tools do have scattered permission systems, like Gmail scopes and Salesforce access levels, but they are not unified enough to provide real control. Composio’s answer is two-layered governance: deterministic access control that decides what the agent can reach, and natural-language policies that constrain what it can do with that access, such as not deleting too many emails or not emailing outside a domain.
The final primitive is reversibility. In code, mistakes can usually be walked back. A bad commit can be reverted, and even if production breaks, there is at least a path to undo the damage. That ability to recover makes it easier to let agents act with some independence.
Knowledge work often has no such undo button. Deleted emails, sent messages, and wire transfers are hard or impossible to reverse. That means the trust model changes: instead of trusting the agent after the fact and fixing errors later, you have to trust before it acts. Composio handles this in two ways. If an action can be reversed, it exposes a reverse path. If it cannot be reversed, the agent must try it in a sandbox first, where the user can review it before anything touches production.
Vaidya closes by saying the bottleneck used to be the model, and for two years everyone raced to make models better. Now the models are good enough that software engineering can be fully autonomous, and the bottleneck has shifted to everything around them. The same models can do hiring, sales, and other knowledge work, but only if the missing infrastructure is built.
He says that is what Composio is building: the layer that gives agents history, context, verification, guardrails, and ways to avoid irreversible mistakes. The company is already powering more than a billion tool calls in total and 300 million tool calls every month. His final message is that the models will keep improving, but the real constraint will increasingly be the substrate around them.
Verdict: The talk is directionally credible on the systems problems around knowledge-work agents, but it mixes solid architectural ideas with promotional overstatements, vivid anecdotes, and self-reported metrics that remain unverified.
Tanmai Gopal opens by naming the central anxiety behind building a “company brain”: if you give an AI system broad access to company knowledge, it may leak secrets. He frames the problem through familiar workplace scenarios — a new trainee seeing things they should not, or an internal assistant exposing sensitive job details — and argues that this fear is what has held back wider deployment of such systems.
He introduces himself as the CEO and co-founder of PromptQL and says the team’s background goes back to Hasura GraphQL, where they worked on data access problems for large customers. That history, he says, gave them both the technical foundation and a “love-hate relationship” with data security.
Over the past year, Gopal says PromptQL has worked closely with a small set of early partners, roughly 15 to 20 so far. Those customers fell into three broad groups: AI-native companies that will move fast if something works, tech-forward companies like Instacart that want the best available tools even if they are imperfect, and highly security-conscious Fortune 100 banks.
He emphasizes that the bank environments are especially important because they cannot afford unreliable AI agents inside sensitive systems. Still, he says PromptQL has made this work there too, and that experience has taught them a lot about how to build something like the “front lobe” of a company’s brain.
Gopal says their version of a company brain is not a single monolithic model. It is, in their case, about 5,000 interconnected pages. Those pages can live as Markdown files in GitHub, as a memory graph, or in some other knowledge structure — the form matters less than the fact that the knowledge is organized and connected.
He then asks what healthy company-brain usage should look like over time. The daily update rate, he argues, should rise as people start using the system more heavily. At first there may be enthusiasm and a burst of setup, but once the system becomes useful, people keep teaching it more: first how to query data, then how to interpret it, then how to act on it, and eventually how to use it for things like A/B tests. In his view, a good company brain does not stay static; its knowledge and update rate both grow.
He draws a sharp line between two use cases. The first is a private assistant use case: an individual asks the company brain to help complete a task, such as answering a client’s security questionnaire by searching internal knowledge. The second is a collaborative one: multiple people use the same agent inside shared tools like Slack to coordinate work, such as incident response.
In the collaborative case, he describes a shared agent helping a team fetch logs, inspect databases, create pull requests, deploy to test and production, and set alerts. Both cases are useful, he says, but both create the same security challenge: how do you let an agent use company knowledge without exposing the wrong information?
Gopal defines a company brain as shared context in Markdown files plus the rules that govern access to data and tools for a software agent. He is careful to distinguish this from the generic idea of a large language model doing everything. The real task, he says, is to build software agents that solve common company problems.
He rejects the idea that the answer is to build one huge knowledge base and then secure it afterward. That, he says, does not work. Nor does it work to imagine some central team building the whole brain for a mature company from scratch; instead, every team member has to own and build their part of it. He describes this as something that should evolve gradually, not be force-built all at once.
To make the problem concrete, he walks through a simple example. Someone receives an email from a customer containing a security survey. They ask the AI agent to search the company brain and help draft answers. The agent reads internal context, pulls the needed facts, writes a response, and the user sends it.
That sounds straightforward, but the hard part is how one person’s knowledge becomes available to another person’s agent. Gopal considers three naive approaches. The first is to have everyone manually write shared knowledge into GitHub, but he says most people will not do that for each other. The second is to build separate “team brains” inside Slack or another siloed workspace, but that just creates another isolated island of knowledge.
His preferred approach is a single company-level wiki made of linked Markdown files. Each file can define who may read or write it, and the important rule is that the agent should not add memory automatically. Instead, it should suggest additions, including which access ranges or domains they belong to, and a person should approve or reject them.
He shows this as a user interface pattern: after a conversation or email reply, the system suggests a set of facts to add to the wiki, and the human can click “Add to Wiki” if the facts are correct. What matters, he says, is not whether the content lands in one file or another; what matters is that the facts are true and the access boundaries are right. Every wiki page should have clear domains — for example, finance or personal — and the system should respect those boundaries.
Gopal then lays down two rules. First, everything should live in one company-level wiki. Second, every change should be attributed to a person, not to the AI agent. He stresses that this traceability matters because if something goes wrong — for instance, if someone accidentally exposes salaries — there must be a named human responsible.
From there, he explains the second piece: scopes. The agent should always read the correct part of the wiki using the current user’s credentials or claims, meaning the identity and permissions of the person asking the question determine what the agent can access. In his model, the agent does not bypass permissions; it behaves more like a human employee with the same access rights.
Gopal then turns to the more powerful collaborative case, where multiple people and an agent work together on the same problem. This, he says, is where the company brain creates the most value. He gives an example from an SRE team dealing with a crashing learning-and-wiki system. One person investigates, learns that a custom prefix is causing problems, and another person joins in to argue through the technical decision. That discussion surfaces the real root cause: a design choice was never documented.
This kind of conversation produces high-quality context because people challenge each other, correct assumptions, and converge on the underlying issue. The problem, he says, is that if the same agent can also deploy to production, the privilege escalation becomes dangerous. The engineer who may be allowed to open a pull request should not automatically be able to use the same agent to push to production.
To make shared AI safe, Gopal says the system should always use the user’s credentials to read context and execute tools. Credentials should not be stored in a sandbox or testing environment. Instead, the system should inject the user’s identity through the HTTP or SQL layer, or through some proxy or virtual environment, so the AI behaves like a person operating within real permissions.
He summarizes the design as a small set of principles: keep all context in one wiki, make every change human-owned, do not let agents write memory automatically, and ensure tools and data access are controlled by user credentials. In his view, if you work backward from those rules, there is really only one logical structure for how the company brain, its permissions, and its tools should fit together.
Gopal closes by saying he would be happy to discuss the architecture further after the talk. He mentions PromptQL again, says they will be launching a new product, and invites people to follow up with him on Twitter and visit the company.
He ends on a lighter note, asking for a photo and telling the audience to stay tuned. He positions their approach as similar to the ideas he has just described: a way to work with company context and AI without being locked into a single cloud vendor, and with the flexibility to use different models as needed.
Verdict: The talk is broadly credible on the high-level security principles of building shared AI systems, but many concrete company/customer claims, implementation details, and success metrics are self-reported and remain unverified from the transcript alone.
Jean-Denis Greze, CTO at Town, begins by saying the talk is not really about Town itself, but about a broader shift in how AI systems work together. He frames the whole topic in terms of search: in most large language model, or LLM, systems, the real job is to make sure the context window—the information the model sees right before it answers or makes a tool call—contains the right data. If the context is right, the model can do its best work.
He traces how this has evolved: first humans manually filled the context window, then retrieval-augmented generation, or RAG, used search tools to pull in information, and now people are building “agentic search,” where the agent has many tools and searches through content to find the right context itself. In his view, the key is still the same: engineer the system so the one important LLM call has the best possible information.
From there, he asks what agent-to-agent really means. He invites the audience to imagine a world with just one agent that has access to all the world’s information: anyone’s email, any company’s data, any government’s records. In that perfect setup, the agent could always produce the best possible outcome. He argues that this is essentially a multi-agent world in its ideal form—one agent with universal access.
But that ideal collides with privacy and security. He brings in the Coase theorem from economics, which says that if everyone had the right information and there were no transaction costs, humans could reach economically optimal outcomes. The problem is that, in reality, we cannot let an LLM see everything. So the question becomes: how well can a system approximate that ideal without violating privacy boundaries?
The first approach he describes is simple: give an agent broad access within a trusted boundary. He uses a family example, where his wife and he share an agent that can access both of their email accounts. In a company, the same idea could apply to an HR team agent that has access to HR systems the way a member of that team would.
This works well because it matches familiar enterprise software models, especially for IT and security teams. But he sees a fundamental weakness: it does not reduce the number of humans involved, and it does not truly break down silos. It just creates a new silo around the person who now has to think through the data.
The second strategy is more selective. Instead of opening full access, you build tools that make a specific trade-off between power and privacy. He gives an example: asking whether anyone in a company is connected to someone on the finance team at Acme Corp. Rather than give full Gmail access to the agent, you could build a tool that scans internal email and returns only a relationship strength score.
The agent would then use that score to find the right employee, contact them through Slack, and request an introduction. He says Town uses this kind of pattern for some features, asking what privacy-preserving tools users would tolerate because they create natural network effects by connecting people across silos. He also mentions another example: letting people draft emails in someone else’s inbox to save time on introductions.
Still, he says this approach is limited because it depends on humans designing the tools and explaining them to everyone. It is manual, and it does not naturally improve just because the models get better.
The third category is shared spaces where data accumulates over time. He likens this to personal wikis, but at the team or company level. Examples include shared skills in a codebase, or a wiki or Airtable-style system where people keep adding useful information. Agents can access these shared silos, so data that is appropriate to share no longer stays trapped in private systems.
He says an especially promising version is a “sweeper AI”: an AI inside each private silo that knows what must remain private, knows the shared spaces available, and at the end of the day moves acceptable new information into public spaces. The key challenge is deciding what can be shared. One option is to ask humans to approve the transfer. Another, which he thinks is coming quickly, is to let the LLM enforce a policy directly.
He predicts that within six months many companies will trust an LLM to surface more private information into shared spaces, especially in smaller, high-trust companies where the sensitive areas are clearer, such as finance and HR. He sees this as a way to improve the trajectory of common work.
The fourth approach is the most familiar: use humans as the bridge. In this model, my agent asks your agent, but the human owner sees and approves the request. Jean-Denis says this is traditional agent-to-agent interaction, with humans acting as the permission layer.
The drawback is scale. If a question requires input from only a few people, the request may get blasted to everyone. In his company-size example, a hundred-person company could end up with a hundred Slack pings asking whether anyone knows a contact at Acme Corp. That is workable, but inefficient.
He then describes what he sees as a more powerful version, even though he has not seen it widely in practice: the black-box approach. In this model, an LLM with full access to the relevant data determines the answer first, without exposing the trace to people along the way. Only after it identifies which specific information must be shared does it ask the relevant human for approval.
In his example, every agent in the company quietly checks whether its owner is connected to the Acme Corp CFO. The black-box system gathers the candidates, determines Bob has the strongest connection, and then asks Bob only whether Jean-Denis may be told that Bob knows the CFO. The trust requirement is high: the company must trust the system to avoid leaking information too early, but he says this can be acceptable in a corporate setting if the human-in-the-loop step is carefully designed.
He also points out a danger: even this model can surface sensitive information indirectly, such as revealing that someone is interviewing elsewhere by asking whether they know a recruiter at another company.
Jean-Denis says he thinks the shared-wiki approach is likely to deliver the quickest return on investment, especially in open-source-style environments and smaller companies. But he also lists the risks. Prompt injection can contaminate a silo if bad content gets pulled into the search process. Shared wikis can go off the rails if an LLM makes one mistake that then gets reused forever. He gives a personal example: his memory system still thinks his agent’s name is Apex, even though he renamed it to Ivy a month ago.
He warns that without human review, false disclosures are inevitable, and those can range from harmless to career-ending or legally damaging. He also notes that black-box systems cannot truly remain black boxes forever; eventually someone in a security or compliance role will want to audit what is happening.
Still, he argues that the direction is clear. Just as coding tools moved from approval mode to “YOLO” and then to auto mode, agent-to-agent systems across information silos will likely move the same way. He expects low-sensitivity data to be shared automatically first, then more and more decisions to be delegated to the model as policies improve and models become more capable. The winning strategy, he says, is to define a low-sensitivity zone where the LLM is allowed to decide, then let that zone expand over time.
In closing, he returns to the network-effects idea. He says the really interesting frontier is not just within one company, but across companies. If multiple organizations could agree to let a common agent work across their information silos, that could unlock something much larger.
He mentions an example from the finance world, where investment banks may benefit from sharing private data about private companies for lending decisions. In that setting, companies are beginning to trust one another’s agents to operate across what used to be separate private systems, with agents deciding what can be accessed. He sees that as a powerful beachhead for the future.
Jean-Denis ends on a note of cautious optimism. He is not sure he personally trusts a future where agents make all privacy decisions, but he believes that is the direction the technology is heading.
Shu Fang from Two Sigma opens with a quick introduction to the company and a playful explanation of its name, then immediately shifts into a legal disclaimer: the views are his, not necessarily the company’s, and he is not endorsing any other firms or products. From there, he frames the talk around a simple but provocative premise: the company has built a cloud-agent system in which agents run under each user’s own identity.
To make that idea stick, he invokes the movie Us. In the film, people have doubles called “tethered,” and when those doubles break loose and cause chaos, they become “untethered.” That becomes the metaphor for the talk: agents should not be loose, anonymous entities. They should be tied to a real user in a controlled way.
He explains that around June 2025, as cloud code and similar tools became more widely used, many agents were still designed to run locally on a user’s machine. That setup worked, but it was limited to the command line interface, or CLI, and it kept the agent on the local device. Two Sigma wanted something broader: agents that could be accessed from mobile, Slack, browsers, and other interfaces, while still running remotely.
The obvious alternative—running an agent as a separate machine identity tied to the user—quickly breaks down, he says. Permissions drift out of sync, licensing becomes harder, some systems do not support multiple identities over the same data, and the boundaries between public and private information get messy. So the team asked a simpler question: why not run the agent as the user directly?
Fang says Two Sigma already had much of the needed infrastructure because the firm had long used remote compute for automation, research notebooks, code containers, and similar workloads. In that setup, every user already had namespaces in Kubernetes clusters across regions, and everything inside those namespaces ran as that user.
He sketches the mechanism at a high level: a controller requests compute, and a sidecar container pulls identity information from a separate service so the actual workload can mount the user’s identity and run as that person. That same pattern, he argues, can support agents just as well as automated jobs.
The biggest internal risk is attribution. If the user and the agent share the same identity, how can the company tell whether a human or an agent took a particular action? Fang jokes that his mustache is there to distinguish him from his agent, but the real need is serious: the firm needs to audit actions, block dangerous behavior, and trace events back to their source.
This is especially important because the company wants to know whether something was done by the human or by the agent acting on the human’s behalf. Without that distinction, the system would be too opaque to trust.
A larger concern is web access. Fang says modern LLM systems need internet access because their models are static and cannot update themselves with current information. Search, fetch, and web tools are therefore core capabilities. But once agents can reach the open web, they also become vulnerable to exfiltration, prompt injection, malware, and accidental use of unlicensed content.
He presents this as the central fear: an agent that can browse freely could leak intellectual property or ingest malicious content. In a regulated finance environment, that risk has to be controlled tightly.
Fang frames the whole problem as a finance-style tradeoff. There is real return in letting agents operate as users, but there is also substantial risk. The goal, he says, is to maximize return while reducing risk, much like optimizing a Sharpe ratio, which measures return relative to risk.
That leads to the two technical goals of the system: first, preserve attribution so actions can be traced to the human or the agent; second, provide safe web access so agents can still use the internet without exposing the firm to unnecessary risk.
The attribution solution is to use a header that every agent appends to as it moves through systems. Fang compares this to trace IDs in observability systems, where a single identifier is propagated through disparate services so engineers can reconstruct what happened.
Here, the same principle applies to agent workflows. By forcing that header to persist through the entire chain, the company can replay the sequence of actions that led to an outcome. That gives better provenance than simply knowing that a particular “shoe agent” or user-triggered process was involved at the start. The actor remains the user, but the path through the system becomes visible.
For safe web access, Fang says the team looked for a way to use search and fetch without letting agents make uncontrolled requests to the public internet. Their answer was Google’s enterprise web grounding offering, which provides access to Google’s web index within the firm’s VPC and network controls. He describes it as offering the same core capabilities they need—search and fetch—while keeping traffic inside the boundary.
The tradeoff is freshness. The index is not fully real time; Fang says it is generally fresh within 24 hours, and for some regularly updated sites within 6 hours. But for most agent use cases, he argues, that is good enough, and it removes the external egress risk.
The second part of the web-access fix is enforcement. Fang says the agent tools for direct web search and fetch are simply denied. Instead of letting the agent use those native tools, the system redirects requests through supported paths such as MCP CLI, client code, or similar harnesses so that all web access goes through the grounded index.
This way, the agent still gets the information it needs, but only through the controlled enterprise path. The point is not to eliminate web use; it is to make web use safe and auditable.
Fang ends the technical portion by summarizing the result: a framework that lets cloud agents run as user identities in a remote environment, with guardrails for attribution and web access. He says the company can deploy a managed fleet of cloud agents for each user, and people who are uncomfortable with CLIs can still interact with agents through other interfaces.
He emphasizes that this is not just experimental. The system was shipped, and employees can already deploy agents in the cloud under their own identities.
During the Q&A, Fang is asked about local LLMs for enterprise use. He says that personally—and not on behalf of the company—he thinks locally managed models may be the eventual direction for much of their token use and inference, for reasons including cost and the volatility of frontier models.
He is also asked about spoofing the header, and he explains that the header alone is not the identity. Someone could populate a header, but that would not let them mimic the underlying user identity, which is still established through the identity system and the chain of RPC entry points. Another question asks how Google’s index helps with prompt injection. Fang answers that the index is internal, cached, and curated for regulated industries, so the risk is reduced because the agent is not freely browsing arbitrary external content.
In one final exchange, he explains that behavioral data and session data help the firm configure agents in a way that improves the user experience while keeping sensitive activity localized. When asked how individuals or teams can create agents across the company, he says the infrastructure is already provisioned for every user. Developers can build agents using existing frameworks, deploy them into their namespace, and then move them through the normal production and security review process if they are meant for broader company use.
He closes by inviting people to talk afterward and noting that the company is hiring.
Verdict: The talk is broadly credible on general enterprise AI/security principles, but many of the concrete implementation, product, and timeline claims are self-reported and remain unverified, and a few security benefits are overstated beyond what the evidence would support.
Ben opens by introducing himself as being from Zo Computer and immediately treating Zo as more than a product: it is “my computer.” He says he even built a site just before the talk—a page with an overview of the other speakers, useful links, and deeper research into their recent thoughts and profiles—which he invites people to scan with a QR code.
From there, he frames the talk as both a tour of Zo and a broader argument about personal agents, especially personal cloud agents, and where he thinks the future is headed.
He gives a quick background on himself: co-founder of Zo, former early employee at Venmo, later at Stripe, and co-founder Rob later went on to build Substack. He then lingers on a small design detail: the old Finder icon, which he says Susan Kare designed to represent a human and a machine in harmony. For Ben, that image captures the goal of AI tools—people should feel at one with their machines, not dominated by them.
He asks who misses the way computers used to feel, when the early internet and first personal computers felt exciting and alive. In contrast, he says today’s digital world feels like a messy sea of apps, sites, and services, increasingly expensive and often bundled with AI features people never asked for.
Ben argues that the problem is deeper than bad product design. He calls it “techno-feudalism,” borrowing feudal language to describe a digital economy where users rent access to SaaS providers, who in turn depend on cloud providers and ultimately on the dominant hardware and infrastructure companies. In this model, users are “peasants”: they are fragmented across tools, locked into systems, and asked to pay for products that monetize attention and data while becoming worse over time.
Zo, he says, exists to challenge that structure by giving people a real home on the internet. The goal is to help ordinary people escape what he describes as a Matrix-like situation and regain ownership of their digital lives.
Ben points to real users to show what Zo does in practice. One example is Charlotte, a private chef and life coach in Los Angeles, who uses Zo to host multiple websites and manage invoices, bookkeeping, scheduling, and notes. He says Zo helps her feel clear, calm, and in control.
He defines Zo as a personal cloud: a home in the cloud that belongs to you. Instead of relying on a pile of separate cloud services, you store your data, AI, and hosted services in one place that you own. For more technical users, he says a personal cloud can span local devices too—his own setup includes a Zo computer, a cloud computer, a Hermes instance, websites, APIs, a laptop, a Mac mini, and a phone. Together, that is his personal cloud.
Zo, he explains, is also a personal server with AI built in. He ties this back to an older, simpler way of publishing websites: uploading files to a server with FTP and pushing changes live. In his view, most people should not have to think about deployment or infrastructure. Zo is meant to make that complexity disappear by putting everything—your data, your AI, and whatever you host—into one place.
He contrasts that simplicity with the confusing reality many people face now: switching between multiple AIs, devices, and SaaS tools. Zo is designed to replace that sprawl with a single home base.
To show that the product is not only for technical people, Ben tells the story of Anthea, an early user who is a free-diving instructor and not technical at all. She had been running her life and business across services like Squarespace and Calendly, but replaced those subscriptions with Zo. He says she canceled those tools, stopped being a “peasant,” and moved her whole operation into one place.
Anthea now hosts retreat websites on custom domains, keeps a personal workspace, and expresses herself online through fairies, mushrooms, and drawings. Ben says this is more than just a website builder: it is a way for regular people to express their whole selves online again, something he feels has been missing since the 1990s.
Ben then explains that the public-facing site is only the surface. When someone shows interest in one of Anthea’s retreats, Zo texts her the person’s number so she can call at the moment of intent. She can ask Zo for a payment link and close the deal right away. He says this has helped her book more revenue than before, to the point that her retreats are now almost too popular.
Behind the scenes, Zo handles her retreat database, notes, accounting, and media. She can upload images from a retreat and have Zo place them on the site, and she is self-hosting everything from one personal server.
Ben emphasizes that Zo has to be simple enough for people like Anthea, his parents, or anyone else who is not technical. He says people keep telling him on LinkedIn and X that their parents are using Zo and “coding up a storm,” and he shares examples of a caterer, a recruiter, and a marketing agency all adopting it in different ways.
At the same time, Zo is meant to be powerful enough for the people in the room. It can host personal agents like Open Claude or Hermes, support tools such as Codex or Gemini, and provide a properly configured Linux virtual machine with root access. Users can SSH in, control Zo through an API or MCP, and build almost anything they want.
Ben moves into a demo, again showing the cloud workspace he calls Zo. He says he made the showcased website just then and can build sites on the fly. The platform supports chatting with many models, including one’s own API keys, and even Claude Code.
He walks through the file system and cloud storage, describing Zo as a kind of Dropbox you own yourself. He mentions automation features that can run AI tasks on a schedule, a large library of built-in integrations and skills, and a built-in browser that can log into sites and let Zo buy things or perform other actions. He admits that using Zo to buy items on Amazon is “dangerous for impulse purchasing.”
Zo can also host arbitrary services over HTTP or TCP, and it includes a personal website space with live-coding-style editing. Ben shows his own Zo space, which he uses for things like a Calendly replacement that lets people book time with him only after Zo reviews them. He makes clear that he prefers this controlled setup to letting people book directly.
Ben closes by widening the lens again. The reason they are building Zo, he says, is to create a better internet in which every person has a real presence online, not just companies. He imagines a world where people interact directly with each other and with company AIs, without middlemen.
He argues that in the near future, most interactions will be with agents, and many of those agents will live in the cloud. That raises the crucial question: whose cloud are they in? He points to Claude Team as an example of a cloud agent—a company-level Claude that anyone in the organization can use—but says the benefit mostly accrues to Anthropic, not the customer. In his view, that is “intelligence feudalism,” where intelligence flows upward to the platform instead of being owned by the user or company.
Zo’s aim, he says, is to change that by letting individuals and companies publish and own their own agents, and let those agents improve through real usage. He says there is already a beta version of this agent-publishing idea, and he ends by inviting people to scan the QR code to talk to him, get AI credits, or sign up. He also reminds people that the speaker overview page is available on his Zo space, then thanks the audience.
The speaker opens by joking about the fatigue of the conference’s familiar AI themes, then immediately reframes the talk. He says he wants to focus not on agentic orchestration itself, but on lessons learned from building “proper generative UX and UI” and on a mental model that can help people think differently about software experiences.
He introduces himself as Gus, a general manager at commercetools who leads product, UX, and engineering for zero-to-one products. He outlines the structure of the talk: first, the problem space; then a product demo; then a discussion of UI protocols; and finally the challenges his team faced and the ways they worked around them.
Gus argues that, for decades, people have adapted to software instead of software adapting to them. He says GPT’s arrival in November 2022 gave everyone a glimpse of true personalization, but most experiences remained static even as AI sped up development. The core question, for him, is why software still looks and feels so fixed.
He describes the burden this creates for users who rely on multiple SaaS applications every day. Each app has its own mental model, its own navigation, and its own way of doing things, so the cognitive load that should have been transferred to the machine stays with the person. In his view, AI makes it possible to shift that burden differently.
To make the point concrete, Gus shows examples of dense, overloaded interfaces and says they are hard to start with even when you know the products. He emphasizes that these applications are the result of a huge amount of work by many teams, but that the complexity accumulates over time. The trade-off has been extensive onboarding, because newcomers need a lot of help just to learn how to use each tool.
His larger argument is that this is the history of software so far: many apps, each with different logic, all demanding separate learning. As the number of tools grows, the average user is forced to keep juggling different navigational models and workflows.
Gus says that last August he and the company founder asked a new question from the lens of artificial intelligence: what foundational shifts could happen if people could change radically the way they interact with software? He notes that commercetools is API-first and has more than 300 APIs, and that this question led to the product he is about to demo.
He then pivots into a quick demonstration, saying that a picture is more tangible than abstract explanation. The demo becomes a way to show how the product generates interfaces from user intent rather than from a fixed, static screen.
In the first part of the demo, Gus enters a query to create a sales report for Q1. He explains that everything on the right side is auto-generated, with the system deciding placement, information architecture, and which components to retrieve from the catalog. But he is blunt: he does not like the result.
He walks through several early variations and keeps finding them confusing. The system keeps changing the meaning of “Q1,” sometimes turning it into January to March, sometimes adding too many KPI cards or charts, sometimes producing layouts that feel inconsistent from one turn to the next. His conclusion is simple: he would not ship that to production.
Gus then shows the current state of the product, which he describes as much more sophisticated. This time, he enters a query about planning a campaign and explains what happens behind the scenes. An orchestrator extracts the user’s intent, locates relevant tools—first-party or third-party—and combines their outputs with context, including agents on MCP servers, to give what he calls enough “ammunition” for a UX agent to render something meaningful.
This newer result feels better to him. He stresses again that AI is guiding the placement and composition, but that the team is still shaping the experience. The system is already live in pre-production, and he uses that as proof that the approach is not just theoretical.
From there, Gus shifts to the different ways UI can be rendered through protocols, depending on how much control the team wants over the experience. He begins with the most controlled approach: the component is shipped as-is, and the agent merely selects and displays it exactly as designed. He says this works well in some businesses, especially when the company wants strong control over the output.
At the other extreme is fully open-ended generation, where the LLM composes the entire experience itself. Gus gives the example of Claude creating an org chart from a simple prompt. It works, but he says he would be reluctant to delegate that much autonomy in a company setting, because the output and outcome would be too hard to control. His UX instincts push him toward maintaining judgment and design authority.
The approach Gus prefers is declarative, which sits between rigid control and total freedom. He says this is the model his team chose. In this setup, the orchestrator classifies intent, invokes tools, retrieves data, maps eligible components from the catalog to the tool outputs, and then broadcasts a UI description—a kind of UI spec. A schema, using Zod for validation, keeps the output compliant, and the final rendering becomes native UI, in their case React components.
He notes that this declarative approach keeps the experience aligned with the design system while still allowing some flexibility. It also avoids one of the problems he saw in the early demos: copy and structure drifting in ways that made the interface feel inconsistent. The system is less deterministic, but not completely uncontrolled.
Gus then turns to the practical challenges. The first one is information architecture: if the agent picks the components, who decides how they are arranged? He says that placement matters a great deal, because random layout would quickly create confusion for users.
To solve that, his team borrows from atomic design, a methodology that organizes interface design into five hierarchical stages. He explains that they taught a UX agent what good looks like by using templates and a structured hierarchy: layout, slots, subslots, and components. In this model, the orchestrator retrieves eligible components, and the system maps them upward through subslots and slots into a template. In effect, the team has codified UX knowledge into the agent.
The second challenge is the design system and component catalog itself. Gus says this becomes the heartbeat of the whole architecture. The catalog is the contract between the agent and the UI, so every property matters. The same applies to layout elements, slots, and subslots, each of which has its own attributes.
He describes this as curation: the careful shaping of the available components, the layout rules, and the protocol compliance needed to make the result meaningful. This is what allows the team to guide the experience rather than merely showcasing a flashy demo. For him, the point is control with purpose.
Finally, Gus says the nature of the work has shifted so much that his teams do not design every pixel anymore. AI can dictate much of the flow, but that change affects PMs and UX designers as much as engineers. The conversation has moved toward schemas, catalog curation, rules, synthetic data generation, query generation, and interaction patterns.
He closes by emphasizing the human side of the transition. If teams want to embark on this journey, they need to pay attention to people, product, and process—the three Ps he uses with other leaders. He leaves the audience with resource recommendations, says the shift is coming, and invites people to connect with him to discuss the topic further.
Jan Curn opens by borrowing the shape of last year’s “MCP isn’t good yet” talk from David Cramer. He recalls how MCP started out clunky and overhyped, but then matured into a widely adopted standard for connecting tools to AI systems. That arc becomes the template for his own argument: x402 is exciting, he says, but still rough around the edges.
He introduces himself as the founder and CEO of Apify, which he describes as a large marketplace for AI tools called “actors.” These actors cover data extraction, automations, and increasingly agentic use cases. Because Apify wants its tools to be easily accessible to agents, he says the company is deeply interested in agentic payments.
Curn argues that agents need money if they are going to do real work without human intervention. In his view, once agents can be trusted with budgets, they can complete longer and more complex tasks. That has led to a flood of competing payment standards from companies across crypto, payments, and tech: he rattles off a long list including L402, MasterCard Agent Pay, x402 from Coinbase, and other proposals from Stripe, Google, Visa, OpenAI, Shopify, Alipay, Apple, OKX, and more.
He frames this as an emerging battleground for the future of agentic commerce. For him, crypto is particularly well suited to the space because traditional payment methods were designed for people, not agents. Credit cards, PayPal, ACH, and bank debit are expensive and poorly suited to microtransactions. More importantly, agent-to-agent interactions raise trust problems: if you do not know who the buyer is, you cannot safely allow chargebacks or disputes. Curn also likes the idea that a decentralized payment rail would not be owned by a single company.
Curn focuses on x402, which he says builds on HTTP status code 402, “payment required,” a code that had sat unused for nearly 30 years. The flow is straightforward: a client makes a request, the server replies with 402, the client signs a payment, a facilitator such as Coinbase verifies it, and the server does the work. Afterward, the payment is settled on-chain.
But he immediately points to a problem: until the payment is actually settled on the blockchain, the buyer can reuse the same wallet funds elsewhere. In other words, the system is vulnerable to double spending. That is acceptable for trivial requests, he says, but not for workloads that take real time or cost real money to complete.
He also notes a standards conflict. x402 wants HTTP 402 as the first response, while MCP-style flows expect HTTP 401. In practice, companies often solve this by creating separate hostnames for each payment flow, like dedicated payment endpoints. Curn thinks that is an awkward anti-pattern, and he argues the protocol should rely more on headers and less on status-code rigidity.
The speaker then turns to billing models. Early x402 payment schemes were designed around exact, fixed payments, which works for simple API calls. Apify’s actors, however, often run for variable lengths of time and use different amounts of resources, so fixed pricing does not fit well.
Coinbase later added an “up to” scheme, where a caller specifies a maximum amount and the server can charge anything up to that limit. Curn says that helps with metered billing in theory, but it still does not solve double spending. One workaround is to charge a fixed amount, do the work, and then refund any unused balance. That works, but it adds a second blockchain transaction and forces the client to trust the server to return the money. He describes that as functional but clunky.
He then highlights a newer approach Coinbase introduced: batch settlement. In this model, the client deposits money into escrow, receives a cryptographic voucher, and uses that voucher to authorize a series of off-chain microtransactions. Those microtransactions are tracked locally rather than written to the blockchain immediately. Later, the system settles them in a batch and eventually releases the remaining escrow.
Curn says this looks promising and that Apify is working on implementing it. He presents it as a better fit for microtransactions and for the economics of agentic work, because it avoids the inefficiency and cost of putting every tiny action on-chain.
Since Apify did not want to create separate endpoints for every payment provider or keep changing its core API, the company built a new service: agi.apify.com. He says “AGI” here stands for “agent general interface,” not artificial general intelligence. It is a simple markdown-based website with instructions for agents on how to buy and use Apify services through x402 or MPP.
The idea is that agents can come to the site, buy a prepaid Apify token, and then use that token through the normal API or through MCP. Because the interface is aimed at agents rather than people, it can change more freely than a traditional API. Curn presents this as a practical way to layer agentic payments on top of existing products without breaking compatibility for thousands of customers.
Curn briefly walks through a demo in which a local wallet is used to allocate $1 through agi.apify.com and receive a 402 payment required response. He mentions that the x402 ecosystem is still so early that Apify had to build its own local wallet tool. The demo does not go smoothly, and he cuts it short with a laugh and says better demos will come later.
He closes by encouraging people to try the technology themselves, saying it is simpler than it may seem and can be set up quickly. The market is still tiny, he says, but he expects it to grow sharply once the era of subsidized tokens ends and agents have to pay real costs for the services they consume. At that point, he believes buying services from external providers will make more economic sense than building everything from scratch, and agentic payments could take off fast.
Verdict: The talk is broadly credible as a product/opinion presentation about early-stage agentic payments, but many of the concrete metrics, launch details, and protocol timelines are self-reported or recent and therefore remain unverified.
Nidhi Kaushik Vyas from Google DeepMind opens by describing multimodal collaborative agents: systems that help users even when their intent is vague, incomplete, or hard to express in the right keywords. She frames the talk around shopping and commerce because those settings make the patterns easier to see, though she says the same ideas apply to other consumer areas like finance and education.
Her core point is that many current agents behave like wrappers around a search bar. They assume the user already knows what they want and can say it clearly. In reality, she argues, people often arrive with only a feeling or a vibe, so the agent has to do more work: uncover preferences, guide exploration, and move the user toward a useful outcome.
She organizes the experience as a loop that starts with fuzzy intent and ends with a successful decision. The first stage is discovery, where the agent gathers context from past conversations, the current query, personal context, and any references the user has shared. From that, it builds a collaborative strategy for what it still needs to learn.
The second stage is research, which she breaks into two parts. First, the agent decides how best to elicit missing preferences. Text alone may not work if the user cannot articulate what they want, so the agent may need visuals, inspiration boards, or other shared references. Then it does the heavy lifting in the background: comparing options, summarizing trade-offs, and returning potential choices.
The final stage is response, where the agent adapts the presentation to the situation. Vyas says many systems fail by replying with too much text. Instead, the agent should choose the right format for the user’s goal, whether that means bullets, tables, or visual boards. She treats response design itself as part of the intelligence.
In the discovery phase, the agent’s job is to remember what matters. Vyas walks through an example where a user wants to redo a living room under a budget. The system gathers session history, user context, and hard constraints stated directly in the query. It also looks for softer constraints, such as style preferences implied by a reference image or a design the user liked.
That is where multimodal input matters. The agent can pull salient signals from images, build a mental model of the user’s taste, and assign confidence to those inferences. It also needs to identify variables that must be refreshed in real time, such as inventory, because stale information would make the recommendation useless.
To evaluate this working state, she says they use auto-raters—automated evaluators that check whether facts are preserved, whether confidence stays within acceptable error bounds, and whether the system is sensitive to counterfactual changes. If a query changes, the extracted constraints should change too, while irrelevant ones remain stable.
Discovery is also about figuring out what the system does not yet know. Vyas calls this the intent gap: the missing variables that need to be resolved before the agent can answer well. In the living-room example, that might include the room width or better confidence in the style preference.
But the agent does not need to ask everything at once. Instead, it should choose the next question that gives the most information gain. If room width determines whether products can fit, then that may be the best next thing to ask because it meaningfully changes the conversation. The goal is to prioritize the most valuable missing piece, not to interrogate the user endlessly.
She says they evaluate this strategy by checking whether the agent identifies the right blockers, avoids over-asking, and asks useful questions. Question utility matters: a good question should move the dialogue forward rather than merely adding friction.
The research phase begins once the system has enough context to ask more targeted questions. Vyas explains that the agent first builds a bridge between a constraint and the product ontology, meaning the structured way the catalog or knowledge base represents products and attributes. That mapping lets the system retrieve relevant items later.
Next, the agent decides how to elicit the remaining preference. If the constraint is subjective—like style—the best method may not be text. Instead, it might show a visual preference board. The agent then chooses examples from the product space that resemble the user’s reference image so it can establish a common language around taste.
She also describes the importance of reading user reactions. Hovering, clicking, and similar micro-signals help the system refine its confidence in what the user prefers. For this stage, they evaluate how efficiently the agent discovers hidden preferences, how many turns it takes, and whether the format matches the nature of the question. Straightforward facts may work well in text, but fuzzy preferences often call for visual anchors.
Once the agent has uncovered the user’s preferences and priorities, it still has to decide how to present the result. Vyas emphasizes that response format should match the intent. If the user wants policy or review details, a summary or bullet list may be best. If they want to compare products, a trade-off table makes more sense. If they want inspiration, then visual references and example images can help them explore possibilities.
She says evaluation at this stage focuses on format accuracy, data fidelity, and actionability. The response should make the relevant information easy to spot, avoid hallucinations, and leave the user confident enough to take the next step, such as making a purchase.
Vyas closes the main talk with four takeaways. First, systems should be designed to accept fuzzy intent, not just cleanly phrased queries. Second, agents should show and ask, not only ask—visuals and comparisons can reveal preferences much faster than text alone. Third, the answer itself must be shaped carefully, because the presentation format is part of the intelligence. Fourth, the loop must be graded with auto-raters at every stage so the system can improve as it grows.
She notes that these auto-raters themselves evolve over time. They may start simple, but they need to grow with the product and its capabilities.
In the Q&A, she is asked how the merchant-side ontology should be structured so that it works well for both agents and merchants. Vyas says merchants contribute domain expertise, and the system depends on them to define how user constraints map onto product metadata. She mentions that they also launched “UCP” to help merchants speak a common language with the agent.
Asked whether merchants or the agent should decide the response format, she says the agent should own that decision. The goal is a horizontal common layer across merchants and a seamless user experience, so the response presentation stays part of the agent’s intelligence.
When someone asks what happens if the user is another agent, she says they are still early in that area, but expects an MCP to serve as the interface. For now, their user studies suggest that people still like being involved, especially in upper-funnel journeys such as discovery and inspiration. She says users seem more likely than agents to drive those early exploration steps, while agents may be more useful lower in the funnel for comparison or negotiation.
Verdict: The talk is broadly credible as a product/design overview of multimodal shopping agents, but many of its most specific performance claims, internal metrics, and “recently launched” platform references remain unverified because they are self-reported and not independently checkable from the transcript.