メインコンテンツへスキップ

詳細を入力してウェビナーの視聴を開始

続行すると、弊社の利用規約、プライバシーポリシーに同意し、データが米国に保存されることに同意したことになります。

このウェビナーを共有する

データとAIのスキルギャップを埋める

私たちは、組織全体にわたってデータとAIのスキルを向上させるために独自に設計された唯一のプラットフォームです。貴社に合わせたプログラムを見てみましょう。

エンタープライズデモを予約する
少人数のチームでの導入をご検討ですか?今すぐ始める
Artificial Intelligence

[RADAR 11x] The New Developer Workflow

October 2026
Webinar Preview

あなたのプレゼンター

Joao Mouraのプロフィール写真

Joao Moura

CEO at CrewAI

のロゴ

Joao Moura is the founder and CEO of CrewAI, a platform that helps companies build and deploy AI agents at scale. He has nearly 20 years of experience in software engineering, product design, and data science. Before starting CrewAI, he worked in engineering and product roles across multiple industries. He studied at NYU Stern School of Business and has partnered with companies like NVIDIA, IBM, and PwC on AI innovation.

Ali Arsanjaniのプロフィール写真

Ali Arsanjani

Director of Applied AI Engineering at Google

のロゴ

Ali Arsanjani is Senior Director of Applied AI Engineering at Google, where he leads the AI/ML Center of Excellence. He also teaches data science at San Jose State University. Before Google, he spent over 20 years at IBM as a Distinguished Engineer and later led machine learning teams at AWS. He holds a PhD in Computer Science and has founded multiple AI-focused ventures.

Fady Mater

Global VP of AI Engineering at DataRobot

のロゴ

Fady leads AI engineering services at DataRobot, helping enterprises build and scale AI and machine learning systems. He brings over 15 years of experience delivering technology solutions across industries, with expertise in distributed systems, enterprise software architecture, and AI-driven digital transformation. Before DataRobot, he held engineering and technology leadership roles at companies including GoDaddy, Rakuten Loyalty, and Spotter.

Kapil Guptaのプロフィール写真

Kapil Gupta

VP of Engineering for Digital Platforms & Enterprise AI at UPS

のロゴ

Kapil leads engineering for digital platforms and enterprise AI at UPS, overseeing large-scale eCommerce systems and digital platform modernization. He specializes in cloud architecture, scalability, and applying AI to enterprise technology. His work at UPS has included modernizing the company's global website, featured at Adobe Summit 2025. Previously, he held senior engineering and technology roles at AT&T and Verizon.

Summary

Software engineering has quietly become AI's proving ground, and two guests who build and oversee agent systems for a living spent this RADAR 11x session laying out what has actually changed versus what just sounds like it has.

Joao Moura, CEO of agent orchestration company CrewAI, and Dr. Ali Arsanjani, senior director of applied AI engineering at Google and a senior lecturer in data science at San Jose State University, joined DataCamp senior software engineer Dan Denney to talk through how AI has reorganized software teams over roughly two years. Moura described a shift from hand-written code to agent-built systems inside his own company, which he says now runs on hundreds of agents. Arsanjani compared today's AI coding tools to early self-driving cars: capable, but not yet trustworthy enough to take your hands off the wheel. Both argued that the real work has moved from writing code to specifying what the code should do, testing it, and building governance around who, or what, is allowed to ship.

They also tackled the economics of running AI-first teams, including token costs, the shrinking line between humans reviewing code and agents reviewing agents, and why unmanaged AI use inside a company risks becoming a more dangerous version of shadow IT. The conversation closes with direct advice for developers trying to keep pace: there's no map for where this goes, so the people who experiment now will be better positioned than those who wait to be told the answer.

Key Takeaways

  • CrewAI's own codebase shifted from being written by hand to being built largely by agents in about two years, according to CEO Joao Moura.
  • AI coding tools have reached a capability level Dr. Ali Arsanjani compares to early supervised self-driving cars: useful and improving fast, but still requiring a human foot near the pedals.
  • Spec-driven, test-driven development is replacing open-ended "vibe coding," using written specifications and markdown files to constrain what an AI agent is allowed to generate.
  • Production AI systems depend on a deep technical stack, spanning data infrastructure, model routing, retrieval and memory components, evaluation tooling, and observability, not a single all-in-one platform.
  • Human-to-agent ratios vary by risk: back-office use cases tolerate more autonomy, while front-office and customer-facing work still runs through human review.
  • A maker-checker pattern, where one system or person generates work and a separate one verifies it, is becoming standard practice for AI-generated output.
  • Companies that don't govern their agents risk a more severe version of shadow IT, with disposable apps, presentations, and artifacts multiplying faster than anyone can track them.
  • More AI capability has translated into more total work for engineering teams, not less, making token efficiency, not just token volume, a real cost consideration.
  • Both guests recommend developers build their own filtering systems, including personal AI agents that track relevant newsletters, blogs, and research, to keep up with the pace of change.

Deep Dives

From Hand-Written Code to Agent-Built Systems

Moura pointed to his own company as the clearest example of how fast the ground has shifted. CrewAI, the agent orchestration platform he runs, was mostly hand-coded when it launched two years ago. That's no longer how it works. "A lot of... how we was written and created two years ago was code written by hand, what is... is not a thing anymore," he said, describing a team of roughly seventy people now supported by hundreds of agents reused across different use cases, interacting daily with a mix of coding agents.

He pointed to a pattern among the companies adopting agents fastest: their underlying data work was already done before they started. Their data lakes were organized, their use cases understood, and components like vector databases or retrieval systems were already shared across projects. Companies that tried to keep every agentic use case under one team's tight control, he said, tended to hit a bottleneck instead of scaling.

To describe where the industry sits today, Moura reached for a car analogy: decades separated the invention of the automobile from the arrival of the seat belt, and AI is at a similarly early, unprotected stage. "We don't have seat belts just yet," he said.

Arsanjani offered a related comparison from his own work evaluating coding agents at Google: today's tools resemble Tesla's early supervised self-driving systems, meaningfully better than a few years ago but not yet hands-off. "It's better than the rest. And I trust it to a large degree, but I keep my vigilance and my foot on the pedals," he said.

Both guests agreed the shift isn't really about typing less code. It's about redefining what counts as engineering work in the first place. "Being an engineer is way more than writing code," Moura said later in the conversation. The work that remains, both argued, is understanding systems well enough to specify, test, and verify what the agents produce, the subject the conversation turns to next.

Spec-Driven Development: The New Guardrails for AI Coding Agents

Arsanjani's central worry isn't that AI writes bad code. It's that people stop being able to explain what the code does. He described sitting in on calls, in both academia and industry, where someone presented AI-generated evaluation results they couldn't actually account for. The failure mode he sees most often isn't a broken model. It's a team that let an agent run on autopilot without checking its work.

His fix borrows from a twenty-year-old idea and extends it one step earlier. Test-driven development, writing the test before the code that passes it, is resurfacing as the backbone of responsible AI coding, but Arsanjani argues it now needs to start with the specification itself, before any test gets written. Rather than relying on the guardrails built into a coding platform by default, he said, teams need to constrain models further for their own domain, their own coding style, and their own use case. At Google, that takes the shape of an internal effort called Project CodeWell: markdown files that steer a model toward an organization's specific standards rather than whatever is generically plausible. "You start putting together not just skills, but these markdown files that guide the model in the direction you need it to be," he said.

The payoff, he said, is real: teams get what amounts to an army of coders working alongside them. But that only works if the specification and testing discipline comes first. Left unspecified, the result is what he calls a hamster wheel, a model generating plausible-looking output in a loop with no one checking whether it actually solves the problem.

Moura's own company is named CrewAI, and Arsanjani used it to make the point stick: agent-generated pieces need to work together deliberately rather than just exist side by side. "If the bits and pieces aren't cohesively flowing into each other, you don't really have a crew, pun intended," he said. Written specifications, in other words, are what turn a collection of agents into an actual team, not just a pile of generated parts.

The Real Stack Behind the AI Hype

Ask what a production AI stack actually contains, and Moura's answer cuts against a common narrative. Some AI labs suggest that a sufficiently capable general model will eventually make individual agents unnecessary. "There is kinda like a narrative coming from... the labs that is saying... AGI is gonna take care of you. You're not gonna have to create all these individual agents," he said. But look under the hood of a system actually running at scale, and a different picture emerges. "It's so much more complex... than kinda like what a lot of people are led to believe," he said.

He mapped out the layers he sees in real deployments. At the base sits the data infrastructure: BigQuery, Redshift, Databricks, or Snowflake, often more than one at once. On top of that come the models themselves, with OpenAI and Anthropic as the obvious choices but a growing number of teams testing alternatives through routers like OpenRouter. Above that sits the deployment layer, plus retrieval components built on vector databases and embedding models, sometimes with fine-tuned models and re-ranking added in. Teams serious about reliability add dedicated evaluation and observability tooling, pointing to platforms such as Galileo and Braintrust as examples of what some companies use to track model performance separately from user-facing monitoring. And at the top sits the interface, which Moura noted doesn't have to look conversational at all. A product can run entirely on an agentic back end while presenting a traditional-looking app to the user.

Every one of those layers, he said, is a lever someone has to deliberately tune. Optimizing a single retrieval pipeline alone involves more decisions than most outside observers assume. That complexity is also why he pushed back on the all-in-one-platform story: in his experience advising companies that have moved agents into production, the stack required to make an AI system reliable keeps growing, not shrinking, as the stakes of getting it wrong go up. That growing complexity, both guests agreed, is exactly why team structure and human oversight matter more now, not less.

Humans, Agents, and the Maker-Checker Pattern

Denney asked both guests for a number: what's a sensible ratio of humans to agents on a team right now? Neither gave one. Arsanjani reframed the question around verification instead of headcount. Picture a ten-step process, he said: six of those steps might run through an agent automatically, but the cross-checking and sanity checks on that output still need a human or a separate system in the loop. "There needs to be a maker checker kind of pattern," he said: something produces the work, and something else, with a real confidence threshold attached, checks it before it ships.

Moura said the right ratio depends heavily on the use case and the risk attached to getting it wrong. Coding, he argued, is one of the areas he's most confident handing to AI, partly because labs have spent years specifically optimizing benchmarks for it. Back-office workflows tolerate more autonomy too, simply because the cost of an error is lower. Front-office and customer-facing work is a different story, where people get far more careful about what they ship.

He gave a concrete example from his own sales process. When a CrewAI salesperson finishes a customer call, agents review the transcript, match it against existing use cases, and draft an architecture proposal along with a visual document a prospect could actually read. Rather than sending that straight to the customer, the system routes it to the salesperson first, by email, for a final pass before anything goes out. "The salesperson just got an email from agent... and that eventually goes online and the customer can see it now," he said.

Arsanjani added that this kind of layered trust has to be earned deliberately. Once a workflow has been run "a hundred times" with consistently good results, he said, teams can reasonably assume the hundred-and-first run will work too, while acknowledging that assumption is probabilistic, not guaranteed. That's the real ratio question, in his view: not how many agents versus humans, but which specific steps have earned the right to run without a person checking them first. Both guests pushed back on reading any of this as a plan to cut headcount. "At the end of the day, your people are actually your most valuable asset," Arsanjani said.

Pizza Tokens, Agent Registries, and the New Shadow IT

The conversation's most memorable image came from an old engineering rule of thumb. The "two-pizza team" standard says a team should stay small enough that two pizzas can feed everyone. Denney asked what happens to that math once agents join the team. Arsanjani's answer: budget for the agents too. "It's me and my agents. They don't have pizza. They have pizza tokens," he said, describing a team as the combination of a person, that person's individual agents for long-running tasks, and a shared group agent that supports a whole division.

Moura pushed the economics further. AI labs, he pointed out, are incentivized by token consumption: the more a customer uses, the better it is for the lab's revenue. Companies buying that capability want the opposite. "You're trying to get a token out of it that is worth it more than what you're paying it for," he said, arguing that prompt design, context selection, and memory management all exist to serve one goal: make every token spent produce more value than it costs.

Both guests pointed to a second risk that gets less attention than the price tag: ungoverned sprawl. Arsanjani called for an organization-wide agent registry, tracking what each agent does, what it consumes, and who's responsible for testing it. "You need an agent registry that has all the agents in the organization, what they're doing, what their inputs outputs are, who's testing them," he said, warning that without one, "it's gonna be the Wild West for agents and for coding."

Moura took the idea further, arguing that unmanaged AI use is a more severe version of a familiar problem. Shadow IT used to require enough people and budget to quietly spin up unauthorized tools. Now one person can generate an app, a deck, or a working prototype in minutes. "You don't have any controls on this, you're in for the new shadow IT," he said, describing teams that end up with hundreds of AI-generated presentations and no way to find the one they actually need. His prescription: organizations need to treat certain things, from agent skills to shared templates, as canonical, or scaling past a certain size becomes nearly impossible. Arsanjani agreed, tying the point back to his earlier framework: governance, he said, isn't just about data access and security. It's what makes the next loop, feedback, evaluation, and the next round of specs, possible at all.


関連

webinar

[RADAR 11x] Don't Waste Your Time on AI Pilots

Experts explain how organizations are moving beyond isolated AI experiments to scalable, production-ready adoption.

webinar

Towards 10x Team Productivity with AI

Richie Cotton, Senior Data Evangelist at DataCamp, demonstrates practicable ways to boost team productivity using generative AI and agentic workflows.

webinar

How To 10x Your Data Team's Productivity With LLM-Assisted Coding

Gunther, the CEO at Waii.ai, explains what technology, talent, and processes you need to reap the benefits of LLL-assisted coding to increase your data teams' productivity dramatically.

webinar

[RADAR 11x] Humans On The Loop

Industry experts explore how human roles are shifting from task execution to supervision and judgment as AI systems become more autonomous.

webinar

Using AI To Increase Your Productivity

Industry experts share real-world examples of how professionals across these fields are using AI to get more done with less effort.

webinar

Progress from Junior Developer to 10x Senior

Ran Aroussi—CEO of Automaze, creator of MUXI, and author of Production-Grade Agentic AI—explores how developers can use AI as a force multiplier on their path from junior to senior and beyond.