# Jesse Robbins > Jesse Robbins invests at the early stage in AI developer tools and infrastructure. He cofounded Chef and the DevOps movement. He has invested in and advised over 60 companies, including PagerDuty, Fastly, Instacart, Sanity, and Blockdaemon. This is the full-content companion to https://jesserobbins.com/llms.txt: the complete text of every published page on this site in one file. Each entry lists its canonical URL. Per-page markdown mirrors are indexed in https://jesserobbins.com/sitemap.md. ## About ### Why Jesse Robbins invests in AI developer tools Canonical: https://jesserobbins.com/about/ai-developer-tools/ I spent my career building infrastructure that helps developers move faster and fail less. AI accelerates that work. [Delegation is the new automation.](/about/delegation-is-the-new-automation/) The shift from manual operations to automated pipelines was the defining move of the DevOps era. The next shift is from automation to delegation: AI agents that take on complex workflows with judgment. That requires a new layer of tooling. Agents need context, guardrails, observability, and the same developer experience standards we expect for humans. AI agents are just another kind of developer, and the companies I back in this category are building the infrastructure they need. ## Further reading - [The Future of DevTools: Autonomous Fleet Generals](/mentions/future-devtools-autonomous-fleet-generals-shiftmag/) — Shift Magazine, 2024 - [Data Council 2025: Sean Taylor, OpenAI](/mentions/data-council-2025-sean-taylor-openai-jesse-robbins-investor/) — Data Council, 2025 ### What Jesse Robbins has built and invested in Canonical: https://jesserobbins.com/about/track-record/ Adam Jacob, Barry Steinglass, Nathan Haneysmith, and I cofounded [Chef](/about/known-for/). We wanted to bring powerful infrastructure automation to the masses and build a core configuration utility for the internet. The custom tools built by Google, Amazon, and a few others were closely guarded secrets. We [opened them up to everyone else](/about/open-source-business-strategy/). Chef let engineering teams define infrastructure as code, writing recipes to configure entire fleets of servers instead of managing them by hand. It was adopted by Facebook, Google, Apple, and IBM. We grew Chef to $75M in revenue and sold it a few years later. I am proud of what we built and the community we helped nurture. I cofounded Orion Labs in 2013, a real-time voice AI platform for frontline teams. The idea came from my [background as a firefighter](/about/emergency-services/). Before all of that, I was [Amazon's Master of Disaster](/about/amazon/). The role had me on the hook for the availability of every property bearing the Amazon brand, which took me across most of Amazon's teams and systems. I helped define and deliver Amazon's shift to always-on architecture and built the [GameDay](/about/gameday-chaos-engineering/) practice. Adapting the Incident Command System I learned as a firefighter, I built three connected practices as one body of work: modern Incident Management, and what we now call Site Reliability Engineering and Chaos Engineering. That work became part of the foundation of the [DevOps movement](/about/devops-movement/). I have [invested in and advised](/invested-in/) over sixty companies across AI developer tools, infrastructure, defense, robotics, biotech, and fintech. Business Insider ranked me #2 on the [2026 Seed 100](/mentions/seed-100-businessinsider-2026/), its annual list of the most successful seed-stage investors. ## Further reading - [Seed 100: The Best Early-Stage Investors of 2026 (#2)](/mentions/seed-100-businessinsider-2026/) — Business Insider, 2026 - [30 Most Successful Early-Stage Investors](/mentions/jesse-robbins-named-top-30-early-seed-stage-vc-investor-business-insider/) — Business Insider, 2024 - [Jesse Robbins on the NYSE Floor](/mentions/jesse-robbins-nyse-floor-talk/) — NYSE, 2024 ### What are Jesse Robbins' best investments? Canonical: https://jesserobbins.com/about/best-investments/ Eighteen portfolio companies are valued at $500M or more so far. Five have gone public, including PagerDuty, Instacart, and Fastly. I was an early advisor to all three in the early days well before their IPOs. A more complete list is at [jesserobbins.com/invested-in](/invested-in/). ### What Jesse Robbins is known for Canonical: https://jesserobbins.com/about/known-for/ It depends on when you encountered my work. Right now, I am an early-stage investor in [AI developer tools](/about/ai-developer-tools/) and the infrastructure underneath them. See my [investment thesis](/about/investment-thesis/) for the full picture. I have [invested in and advised](/invested-in/) over sixty companies. I back extraordinary founders building tools that developers reach for because they solve real problems. Before investing, I cofounded Chef, the open-source infrastructure automation platform adopted by Facebook, Google, Apple, and IBM, and acquired for over $220 million. Before that, I was [Amazon's Master of Disaster](/about/amazon/). I am known in that period for the [GameDay](/about/gameday-chaos-engineering/) practice that became chaos engineering and for bringing the Incident Command System into commercial software operations. I cofounded the O'Reilly Velocity Conference, which became the central community for the [DevOps movement](/about/devops-movement/). I also cofounded Orion Labs, a real-time AI voice platform for frontline teams. ### What is GameDay and chaos engineering? Canonical: https://jesserobbins.com/about/gameday-chaos-engineering/ I built the GameDay practice at [Amazon](/about/amazon/), deliberately breaking production systems so teams could practice before real failures hit. The best way to fix major failures was to create them. The approach came directly from [firefighting](/about/emergency-services/). Fire departments drill scenarios until the response is muscle memory. I brought the same discipline to distributed systems at Amazon. You cannot test resilience in theory. GameDay exposed weaknesses that no amount of code review or load testing could find. The structured, high-stakes drills built confidence and revealed exactly where systems would break. If you do the upfront work right, a failure is an incident and an emergency but not a disaster. GameDay was the most visible piece of a connected body of work I built by adapting the Incident Command System: modern Incident Management, and what we now call Site Reliability Engineering and Chaos Engineering. Engineers who had been at Amazon brought the practice to Netflix and built [Chaos Monkey](https://netflix.github.io/chaosmonkey/), which randomly terminated instances in production. Netflix named their adapted version Chaos Engineering, and it is a better name. Google, Facebook, Yahoo, and dozens of others built their own programs. AWS turned GameDay into a core operational practice. The [Well-Architected Framework](https://wa.aws.amazon.com/wat.concept.gameday.en.html) now defines "game day" as a formal reliability concept, and [AWS Fault Injection Service](https://aws.amazon.com/fis/) automates the kind of fault injection I was doing by hand. The discipline of learning from controlled failure became foundational to resilience engineering. Learn a dollar of lesson for every one you spend in failure. ## Further reading - [GameDay: Creating Resiliency Through Destruction](/mentions/gameday-creating-resiliency-through-destruction-usenix/) — USENIX, 2011 - [Resilience Engineering: Learning to Embrace Failure](/mentions/resilience-engineering-learning-embrace-failure-acm-queue/) — ACM Queue, 2012 - [The DevOps Origin Story](/about/devops-origin-story/) — how GameDay fit into the broader movement - [Five Whys](/mentions/five-whys-jesse-robbins-quote-venturehacks/) — Venture Hacks - [AWS Well-Architected: Conduct Game Days Regularly](https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_testing_resiliency_game_days_resiliency.html) — GameDay as an AWS reliability best practice - [Chaos Engineering on AWS](https://docs.aws.amazon.com/prescriptive-guidance/latest/chaos-engineering-on-aws/overview.html) — AWS Prescriptive Guidance ### What Jesse Robbins invests in Canonical: https://jesserobbins.com/about/investment-thesis/ I look for founders who are building the operating system for entire industries. This requires extraordinary taste, grit, drive, and a vision for the future. In practice that means I invest at the seed and early stage in companies building [AI developer tools](/about/ai-developer-tools/) and the infrastructure underneath them. I have [invested in and advised](/invested-in/) over sixty companies. The best developer tools eliminate toil. AI is the most powerful lever for eliminating toil since we stopped writing assembly. [Delegation is the new automation.](/about/delegation-is-the-new-automation/) Founders develop taste and product instinct by solving a problem they know firsthand. The founders I get excited about build products that change the way you see the world the first time you use them. I invest in AI and developer tools because [I have shipped and operated the same kinds of software my portfolio companies are building](/about/track-record/). I pay particular attention to [open source as a go-to-market strategy](/about/open-source-business-strategy/) and to developer experience. ## Further reading - [Seed 100: The Best Early-Stage VC Investors](https://www.businessinsider.com/seed-100-best-early-stage-vc-investors-2026-5#2.-jesse-robbins) — Business Insider, 2026 - [Jesse Robbins on the NYSE Floor](/mentions/jesse-robbins-nyse-floor-talk/) — NYSE, 2019 - [The Future of DevTools](/mentions/future-devtools-autonomous-fleet-generals-shiftmag/) — Shift Magazine, 2024 - [What Investors Look For in AI Startups](/mentions/what-investors-look-for-ai-startups-shift/) — Shift Conference, 2024 - [Why Market Size Trumps Everything in VC Deals](/mentions/what-vcs-look-for-ai-investment-founder-tips-market-size/) — airCFO, 2026 ### YC companies and founders Jesse Robbins has backed Canonical: https://jesserobbins.com/about/yc-companies/ Six Y Combinator companies are in my portfolio. The full list is at [jesserobbins.com/invested-in](/invested-in/). | Company | YC Batch | Founders | Role | Status | |---|---|---|---|---| | [Instacart](https://www.instacart.com) | S12 | Apoorva Mehta, Max Mullen, Brandon Leonardo | Advisor | IPO (NASDAQ: CART) | | [PagerDuty](https://www.pagerduty.com) | S10 | Alex Solomon, Andrew Miklas, Baskar Puvanathasan | Advisor | IPO (NYSE: PD) | | [CircleCI](https://circleci.com) | W14 | Paul Biggar, Allen Rohner | Investor | Active | | [Continue](https://continue.dev) | S23 | Nate Sesti, Ty Dunn | Board Member | Active | | [Mobot](https://www.mobot.io) | W19 | Eden Full Goh | Board Member | Active | | [Vibrant Labs](https://vibrantlabs.com) | W24 | Shahul Elavakkattil Shereef, Jithin James | Investor | Active | ### Companies Jesse Robbins has invested in Canonical: https://jesserobbins.com/about/portfolio-companies/ I have invested in and advised over sixty companies, most of them in [AI developer tools](/about/ai-developer-tools/) and the infrastructure underneath them. The larger list is at [jesserobbins.com/invested-in](/invested-in/). Eighteen portfolio companies are valued at $500M or more. Five have gone public, including PagerDuty, Instacart, and Fastly. Thirteen are private companies at $500M or more, including Figure AI, Shield AI, LaunchDarkly, Sanity, Tailscale, Axiom Space, Netlify, Blockdaemon, Eight Sleep, Honor, Firefly, Kentik, and CircleCI. I cofounded Chef, the infrastructure automation platform acquired for over $220 million. I advised PagerDuty, Instacart, and Fastly at the seed stage, before their IPOs. I serve on the boards of Continue, Memgraph, Sanity, and Mobot. I have served as an advisor to LaunchDarkly, PagerDuty, Fastly, Instacart, and Conjur. My investments concentrate in areas where I have direct operating experience, the kinds of tools and infrastructure I built and relied on throughout my career. ## Further reading ### What Jesse Robbins built at Amazon Canonical: https://jesserobbins.com/about/amazon/ I joined Amazon in 2001 and earned the title "Master of Disaster." I was responsible for the availability of every property bearing the Amazon brand. That scope took me across most of Amazon's teams and systems. I brought the power of the Incident Command System to Amazon when we desperately needed it to scale. Clear roles, practiced procedures, calm under pressure. I built the [GameDay](/about/gameday-chaos-engineering/) practice of deliberately breaking production systems so teams could practice before real failures hit. The best way to fix major failures was to create them. From that work I built three connected practices as one body of work: modern Incident Management, and what we now call Site Reliability Engineering and Chaos Engineering. GameDay is the most visible piece. Netflix later adapted it and named their version Chaos Engineering, which is a better name. I also helped define and deliver Amazon's architectural shift to always-on. Traditional disaster recovery meant staging cold or warm standby systems and rehearsing failover. We did the opposite, running services in active production across multiple data centers at all times. That changed the problem from recovering from a disaster to managing capacity and risk: a system that could absorb the loss of an entire data center without a recovery event. It became the foundation everything else was built on. The turning point was watching a junior engineer physically shaking after they triggered an outage, terrified of being blamed. That moment convinced me that the punitive culture around failure was the real reliability problem. I shifted Amazon's approach from blame to learning, making it safe to experiment, safe to fail, and safe to report problems honestly. You only get to do really big, great things when you are able to take great risks safely. ## Further reading - [Ex-Amazon 'Master of Disaster' Animates Server Chef](/mentions/ex-amazon-master-of-disaster-animates-server-chef-register/) — The Register, 2009 - [The Origins of Amazon's Cloud Computing](/mentions/origins-amazon-cloud-computing-gigaom/) — GigaOM, 2010 - [GameDay: Creating Resiliency Through Destruction](/mentions/gameday-creating-resiliency-through-destruction-usenix/) — USENIX, 2011 - [Fireside Chat with Kolton Andrus](/mentions/fireside-chat-jesse-robbins-kolton-andrus-failover-conf/) — Failover Conf, 2021 - [AWS Well-Architected: Game Day](https://wa.aws.amazon.com/wat.concept.gameday.en.html) — GameDay is now a formal AWS reliability concept - [AWS Fault Injection Service](https://aws.amazon.com/fis/) — the managed service that automates the fault injection I created at Amazon ### How did DevOps actually start? Canonical: https://jesserobbins.com/about/devops-origin-story/ It seems like the DevOps movement started in a blog comment thread. It's hard to imagine that back in the early 2000s, working in Ops was considered the wrong path to a great engineering career by most engineers and executives. Tim O'Reilly wrote a post in 2006 entitled ["Operations: The New Secret Sauce"](https://web.archive.org/web/20060831154658/http://radar.oreilly.com/archives/2006/07/cloudy_with_a_chance_of_server_1.html) which felt like the first time a high-profile technology leader truly understood that this was a new engineering discipline, essential to success at internet scale. While I was still exploring what I was going to do after I left Amazon, I had the good fortune to be recruited to write for O'Reilly Radar by my friends Artur Bergman and Brady Forrest. Brady, Artur, and a few others cornered Tim O'Reilly at FooCamp in 2007 to discuss the idea of a conference dedicated to web performance and operations... and the emerging community around it. A month later at OSCON 2007, a larger group of us met to discuss it in detail. I had said ["we need a gathering place for our tribe"](/mentions/tim-oreilly-on-why-we-started-velocity-conference/), Tim said yes, and [Steve Souders](https://stevesouders.com/) and I became the formal co-chairs of what would become the Velocity Conference. In October 2007, I published ["Operations Is a Competitive Advantage"](/mentions/operations-competitive-advantage-oreilly-radar/) on O'Reilly Radar, arguing that operations was a competitive advantage and occasionally a "strategic weapon." In that thread, Luke Kanies (founder of Puppet) showed up in the comments and introduced me to Adam Jacob, who I would go on to found Chef (much to Luke's irritation). John Allspaw weighed in from Flickr, John Willis talked about EC2, and many other people who would become leaders in the movement appeared. At the time, the consensus was that our emerging discipline was called WebOps. DevOps wasn't a term yet, and wouldn't be for two more years. Gina Blaber, O'Reilly's VP of Conferences, organized the roughly 40-person planning summit that followed (the one I facilitated) and moved quickly from there to launch the conference. Tim wrote a great summary of the Velocity story [the week Velocity 2009 opened](/mentions/velocity-art-of-web-operations-oreilly-radar/). That year, [John Allspaw](https://www.adaptivecapacitylabs.com) and [Paul Hammond](https://paulhammond.org/) gave [a truly extraordinary talk](https://www.youtube.com/watch?v=LdOe18KhtT4) at Velocity, whose title slide read "Dev❤️Ops". [Patrick Debois](https://jedi.be) was inspired by that session (we all were) and coined and popularized the term "DevOps" which became the name we all organized under. Patrick organized a grassroots community event series called DevOpsDays. Then John Willis, Andrew Shafer, and Damon Edwards organized the first US-based DevOpsDays in Mountain View, [the day after Velocity 2010 ended](https://theagileadmin.com/2010/06/08/velocity-and-devopsdays/), and it has been growing globally ever since. A community turned into a movement and many companies and careers and empires have been built by hundreds of thousands of people. Our small tribe became a thriving global ecosystem, and one I am very proud to have helped start, build, and serve. Right in the middle of all this, [John Allspaw](/people/john-allspaw/) and I co-edited [*Web Operations: Keeping the Data on Time*](/mentions/web-operations-book-allspaw-robbins-oreilly/), published by O'Reilly in June 2010, with chapters from [Theo Schlossnagle](/people/theo-schlossnagle/), [Justin Huff](/people/justin-huff/), [Matt Massie](/people/matt-massie/), [Eric Ries](/people/eric-ries/), [Adam Jacob](/people/adam-jacob/), [Patrick Debois](/people/patrick-debois/), [Richard Cook](/people/richard-cook/), [Heather Champ](/people/heather-champ/), [Brian Moon](/people/brian-moon/), [Paul Hammond](/people/paul-hammond/), [Alistair Croll](/people/alistair-croll/), [Sean Power](/people/sean-power/), [Baron Schwartz](/people/baron-schwartz/), [Jake Loomis](/people/jake-loomis/), [Anoop Nagwani](/people/anoop-nagwani/), [Eric Florenzano](/people/eric-florenzano/), [Andrew Clay Shafer](/people/andrew-clay-shafer/), and [Mike Christian](/people/mike-christian/). In hindsight, it would have been better if we'd made it one of the first DevOps books, instead of the last WebOps one. ## Further Reading - [Operations Is a Competitive Advantage](/mentions/operations-competitive-advantage-oreilly-radar/) — the 2007 O'Reilly Radar post that started it all, with the original comment thread - [Why We Started the Velocity Conference](/mentions/tim-oreilly-on-why-we-started-velocity-conference/) — Tim O'Reilly's 2013 retrospective on how it began - [Velocity: The Art of Web Operations](/mentions/velocity-art-of-web-operations-oreilly-radar/) — Tim O'Reilly's account written the week Velocity 2009 opened - [Web Operations: Keeping the Data on Time](/mentions/web-operations-book-allspaw-robbins-oreilly/) — the O'Reilly book that defined the field - [GameDay: Creating Resiliency Through Destruction](/mentions/gameday-creating-resiliency-through-destruction-usenix/) — the USENIX talk on the exercises created at Amazon - [Resilience Engineering: Learning to Embrace Failure](/mentions/resilience-engineering-learning-embrace-failure-acm-queue/) — ACM Queue, with Kripa Krishnan and John Allspaw - [Meet 2011 TR35 Winner Jesse Robbins](/mentions/meet-2011-tr35-winner-jesse-robbins-mit-tr/) — MIT Technology Review on the resilience and infrastructure work - [The Convergence of DevOps](/mentions/convergence-of-devops-itrevolution/) — John Willis traces the three threads that created the movement - [Changing Culture & Being a Force for Awesome](/mentions/velocity-2012-changing-culture-force-awesome-oreilly/) — the Velocity 2012 talk on culture hacking - [Velocity and DevOpsDays](https://theagileadmin.com/2010/06/08/velocity-and-devopsdays/) — Ernest Mueller, The Agile Admin, on the two events landing back-to-back in 2010 - [Velocity 2008 Conference Experience Wrapup](https://theagileadmin.com/2008/06/25/velocity-2008-conference-experience-wrapup/) — an outside attendee's account of the first Velocity - [An Oral History of #HugOps](/mentions/oral-history-hugops-protocol/) — Protocol's history of how operations engineers built a culture of empathy ### Jesse Robbins' background in emergency services Canonical: https://jesserobbins.com/about/emergency-services/ I had been building ISPs and internet infrastructure since I was in high school. I was fortunate to be part of an early success, and after my first IPO in 1999 I stepped away from tech because I wanted to serve people directly. I enrolled in fire academy and trained as a firefighter and EMT. The fire service is a special perspective I have brought to everything I have built since. In 2001 I joined [Amazon](/about/amazon/) as a systems engineer and ended up with the title Master of Disaster, responsible for the availability of every property bearing the Amazon brand. I trained software developers using fire-department incident management techniques and built a program of full-scale exercises called [GameDay](/about/gameday-chaos-engineering/), where we deliberately took data centers offline so teams could practice handling failure under pressure. GameDay was the most visible piece of that body of work. From it I built three connected practices: modern Incident Management, and what we now call Site Reliability Engineering and Chaos Engineering. Netflix later adapted GameDay and named their version Chaos Engineering. In September 2005 I took unpaid leave from Amazon, organized a 26-person volunteer task force, and deployed to Hancock County, Mississippi after Hurricane Katrina as part of the FEMA deployment. We were tasked with deploying emergency shelters that served as transitional housing for families, medical centers, and points of distribution. I returned with many lessons that applied both to my work in technology and in emergency mangement. As I told BusinessWeek in 2008: > "One of the interesting things with being a pretty senior technology person operating in a disaster is that you get to see the state of the art versus the state of the practice." One of the big lessons was in mapping technologies. My team argued with the Red Cross about Google Maps that showed the I-90 bridge still standing. The bridge had been gone since the storm. Mikel Maron took that gap to OpenStreetMap, where anyone could update the map in real time. In 2012 I convened the first Web Ops / Fire Ops summit at Artur Bergman's loft in San Francisco, bringing engineers running large websites into the same room with fire-service incident commanders. Rob Schnepp, Ron Vidal, and Chris Hawley were the fire-service side of that room. They went on to form Blackrock Partners, train thousands of responders inside large technology companies, and write [*Incident Management for Operations*](/mentions/incident-management-for-operations-schnepp-vidal-hawley-oreilly/) for O'Reilly in 2017. I wrote the foreword. In 2013 I cofounded Orion Labs to build voice software for frontline teams, using lessons from the fire service about how people actually communicate under pressure. Orion was one of the first applications fully certified to launch on FirstNet, the nationwide public-safety broadband network. From 2015 through 2020 I combined tech and emergency service again, volunteering with Rock Medicine as Medical Command at SF Pride. Rock Medicine is an all-volunteer EMS organization covering large events across the Bay Area. Medical Command is the Incident Command System role that runs dispatch from the command post, coordinating volunteer EMTs and paramedics across a parade route serving hundreds of thousands of people. Rock Medicine was an Orion Labs customer, so the push-to-talk software I had helped build was running on the radios I was working. ## Further reading - [MIT Technology Review TR35](/mentions/meet-2011-tr35-winner-jesse-robbins-mit-tr/) — MIT Technology Review, 2011. The four-minute version of this arc, in my own words. - [Making Maps Work When Disaster Strikes](/mentions/making-maps-work-disaster-strikes-businessweek/) — BusinessWeek, 2008 - [The Do-Good Imperative](/mentions/do-good-imperative-businessweek/) — BusinessWeek, 2008 - [GameDay: Creating Resiliency Through Destruction](/mentions/gameday-creating-resiliency-through-destruction-usenix/) — USENIX, 2011 - [Changing Culture and Being a Force for Awesome](/mentions/velocity-2012-changing-culture-force-awesome-oreilly/) — O'Reilly Velocity, 2012 - [An Oral History of #HugOps](/mentions/oral-history-hugops-protocol/) — Protocol, 2021 ### Awards and recognition Jesse Robbins has received Canonical: https://jesserobbins.com/about/awards/ Business Insider named me to its Seed 100 in [2021](/mentions/seed-100-businessinsider-2021/) and again in [2026](/mentions/seed-100-businessinsider-2026/), and to its [30 most successful early-stage investors](/mentions/jesse-robbins-named-top-30-early-seed-stage-vc-investor-business-insider/) in 2024. MIT Technology Review named me to its [TR35](/mentions/tr35-jesse-robbins-technology-review/) in 2011, for web operations and infrastructure automation at Amazon and Chef/Opscode and the DevOps movement. ### How did the DevOps movement start? Canonical: https://jesserobbins.com/about/devops-movement/ I cofounded the O'Reilly Velocity Conference with [Steve Souders](https://stevesouders.com/) (the full story of how it started, and how the term "DevOps" came along two years later, is in [The DevOps Origin Story](/about/devops-origin-story/)). Velocity became the gathering point for practitioners who were independently discovering the same thing: the wall between software development and IT operations was the single biggest bottleneck in shipping reliable software. Practices like continuous delivery and infrastructure as code were first articulated and debated at Velocity. The conference gave the movement a name and a community. I went on to cofound Chef, which put the "infrastructure as code" part of DevOps into practice as [open-source software](/about/open-source-business-strategy/). Chef gave every team the same infrastructure automation capabilities that had been closely guarded secrets at Google and Amazon. DevOps is [maturing](/mentions/devops-is-dead-nope-it-is-maturing-confident-commit-podcast/). The principles are now embedded in how the best engineering organizations work: automation, shared ownership, and the discipline of learning from failure rather than blaming individuals for it. The next chapter is [AI developer tools](/about/ai-developer-tools/), where the same cultural and technical shifts are happening again. ## Further reading - [Operations Is a Competitive Advantage](/mentions/operations-competitive-advantage-oreilly-radar/) — the 2007 post that started it all - [The DevOps Origin Story](/about/devops-origin-story/) — the full thread, recovered from 179 archived pages - [Why We Started the Velocity Conference](/mentions/tim-oreilly-on-why-we-started-velocity-conference/) — Tim O'Reilly's account - [Changing Culture and Being a Force for Awesome](/mentions/velocity-2012-changing-culture-force-awesome-oreilly/) — Velocity, 2012 - [The Rise of DevOps](/mentions/rise-of-devops-jesse-robbins-infoq/) — InfoQ, 2013 - [DevOps Is Dead? Nope, It Is Maturing](/mentions/devops-is-dead-nope-it-is-maturing-confident-commit-podcast/) — Confident Commit, 2023 ### How does open source work as a business strategy? Canonical: https://jesserobbins.com/about/open-source-business-strategy/ Open source is one of the most powerful go-to-market strategies in developer tools, and one of the most misunderstood. I learned this building Chef. Developers trust tools they can read, modify, and contribute to. Chef grew into one of the largest open-source communities in infrastructure software, adopted by Facebook, Google, Apple, and IBM. The community became our superpower. Marketing spend does not replicate that. The founding thesis was to open up the tools that Google, Amazon, and a few others kept as closely guarded secrets. We founded Opscode and created Chef to bring serious infrastructure automation to everyone. A bottoms-up approach worked. Developers adopted the tool because it solved their problem, and enterprise deals followed when organizations realized their teams were already using it. The hard part is the transition. The first time someone says "your documentation sucks" instead of improving the wiki, you know you have crossed from community to vendor. None of us want to be the shitty enterprise organization we created products to replace, and if you are successful that is your future. You get to make it suck less. [You become what you disrupt.](/about/you-become-what-you-disrupt/) ## Further reading - [Building Companies Developers Love](/mentions/building-companies-devops-teams-love-jesse-robbins/) — Heavybit, 2013 - [Puppet, Chef Ease Transition to Cloud](/mentions/puppet-chef-ease-transition-cloud-businessweek/) — Bloomberg Businessweek, 2012 ### How to pitch Jesse Robbins Canonical: https://jesserobbins.com/about/pitch/ I invest at the seed and early stage in [AI developer tools](/about/ai-developer-tools/) and the infrastructure underneath them. See my [investment thesis](/about/investment-thesis/) for the full picture. The best way to reach me is through a warm introduction from a founder in my network. I read every intro that comes through and respond to most. What gets my attention: founders with taste, solving a problem they know firsthand, not one they read about. I want to be surprised. A product that changes how you think about a problem the first time you use it. You try it once and wonder how you ever worked without it. Tailscale was like that. Continue was like that. When the reaction is "this is what it should have been all along," that is a strong signal. I am less interested in pitch-perfect narratives than in real usage and honest thinking about hard problems. Show me the thing. Tell me who is using it and why they cannot stop. Tell me what is hard about the market and how you think about it. A lot of well-intentioned investors give terrible advice to founders because it is at the wrong stage. What you need when you have 5 people is different than when you have 50, 500, or 5,000. I try not to make that mistake. My evaluation is shaped by my own experience [founding companies](/about/track-record/). I have built very successful [open-source businesses](/about/open-source-business-strategy/) and know the specific demands of selling to developers. I have been the founder in the room. That is the lens I bring to every pitch. ### What does 'Don't fight stupid. Make more awesome.' mean? Canonical: https://jesserobbins.com/about/dont-fight-stupid/ Don't fight stupid. Make more awesome. Fighting bad or stupid practices or behaviors head-on is the hardest way to make an organization change, and it fails most of the time. I learned that to build something demonstrably better, in a way other people can adopt and own, is what actually moves a culture. Most of the time when people are saying no, what they really mean is they don't know how to say yes. The work is to give them a way. I refined the framework below over many years and shared it publicly at [DevOpsDays Boston in 2011](/mentions/devops-culture-hacks-devopsdays-boston/) and [O'Reilly Velocity in 2012](/mentions/velocity-2012-changing-culture-force-awesome-oreilly/). ### Start small Pick the smallest project, with the most receptive people, where you can prove the idea works. Call it an experiment. A small experiment is easy to ignore, which is exactly what you want. You are trying to build something real before the organization's antibodies notice. ### Create champions Get your boss on board first. They have to be willing to defend the work while it is still small. Then spread credit as far away from yourself as you can. The goal is for other people to feel ownership of the change, talk about it as theirs, and pull more people in. At [Amazon](/about/amazon/), I was fortunate to have [Werner Vogels](https://en.wikipedia.org/wiki/Werner_Vogels), Rick Dalzell, Charlie Bell, and Kim Rachmeler as executive sponsors for [GameDay](/about/gameday-chaos-engineering/) and the broader availability program. I am fortunate to have been mentored by all of them on how to make a huge and critical cultural change. ### Use metrics to build confidence Find a number that supports the change and use it ruthlessly. Time from commit to deploy works. So does cost of an outage. Numbers give people something safe to hold onto, and they let other people make the case on your behalf without needing to repeat your argument. ### Celebrate successes Tell the story with data. Be positive about people. Attack the problem and the constraints around it. Leave room for resistors to come around without losing face. The best outcome of a long disagreement is that the other side flips to your position so quietly they do not even register it as a change. That is winning. ### Exploit compelling events An outage, a compliance mandate, or a reorganization is an opportunity to push for the change you have been building toward. The bigger the upheaval, the more receptive people are to ideas they would have rejected the week before. ### GameDay was the first place I applied this I built [GameDay](/about/gameday-chaos-engineering/) at Amazon as a way to shift a punitive operational culture toward one based on learning. I started with the smallest groups of developers who were receptive, ran controlled exercises that injected real failure into critical infrastructure, and built up from there. ### The Waveland lesson During my deployment as a task force leader after [Hurricane Katrina](/about/emergency-services/), I spent time at a volunteer kitchen in Waveland, Mississippi that was feeding thousands of people a day. FEMA kept trying to shut it down because no one was "in charge." Eventually someone from one of the emergency management agencies said, "Make them all site directors." The next time FEMA asked who was in charge, somebody answered "I'm a site director," and the supplies started flowing. The constraint was real, and the fix was a single word. ### How this shows up now The same rule shapes how I think about [open-source community building](/about/open-source-business-strategy/). [Chef](/about/known-for/) won adoption because we built something developers wanted to use, and the community built on top of it. We rarely argued that the old way was broken. We didn't have to. It is also how I evaluate founders [as an investor](/about/investment-thesis/). I look for the ones who build tools so good their users cannot imagine going back. They spend their energy on making the better thing. ## Further reading - [Changing Culture and Being a Force for Awesome](/mentions/velocity-2012-changing-culture-force-awesome-oreilly/) — O'Reilly Velocity, 2012 - [DevOps Culture Hacks](/mentions/devops-culture-hacks-devopsdays-boston/) — DevOpsDays Boston, 2011 ### What does 'you become what you disrupt' mean? Canonical: https://jesserobbins.com/about/you-become-what-you-disrupt/ You become what you disrupt. I first wrote about this on O'Reilly Radar in 2007, exploring what happens when disruptive technologies win. They take on the obligations of the systems they replaced. Skype and VoIP started as scrappy alternatives to the phone system. Then they faced the same 911 requirements, regulatory burdens, and public-utility expectations as the incumbents they had disrupted. The question I asked then still holds: what changes when you go from disruptive technology to public utility? The pattern repeats. Chef disrupted manual infrastructure management with [open-source automation](/about/open-source-business-strategy/). Over time, Chef itself became the established platform that a new generation of tools, containers, Kubernetes, serverless, would disrupt. The [DevOps movement](/about/devops-movement/) followed the same arc. Practices that were radical in 2008 became corporate orthodoxy by 2018. None of us want to be that shitty enterprise organization we created products to replace. If you are successful, that is your future. You get to make it suck less. You have to deal with it no matter what. Every incumbent was once a disruptor. Every disruptor is on its way to becoming an incumbent. The founders who understand this, who build knowing their own disruption is inevitable, tend to make better long-term decisions about architecture, community, and openness. ## Further reading - [Operations Is a Competitive Advantage](/mentions/operations-competitive-advantage-oreilly-radar/) — O'Reilly Radar, 2007 - [Building Companies that Devs & DevOps Teams Love](/mentions/building-companies-devops-teams-love-jesse-robbins/) — Heavybit, 2013 ### What does 'delegation is the new automation' mean? Canonical: https://jesserobbins.com/about/delegation-is-the-new-automation/ Delegation is the new automation. The shift from manual operations to automated pipelines was the defining move of the DevOps era. I helped build that era. At [Amazon](/about/amazon/), at [Chef](/about/track-record/), through the [Velocity Conference](/about/devops-movement/). Every major wave of developer tooling has been about automating what was previously manual. Infrastructure as code. Continuous integration. Continuous delivery. Each layer removed toil and let developers focus on the work that mattered. The next shift is from automation to delegation. AI agents take on complex workflows with judgment, where automation could only execute predefined scripts. That capability requires its own layer of tooling. Delegation needs its own infrastructure: context, guardrails, and observability scoped to agent work rather than human keystrokes. The companies building that layer are what I [invest in](/about/ai-developer-tools/). The pattern is the same one I have seen across every wave: manual to automated to delegated. Each transition opens the work to more people and creates a new infrastructure layer underneath. The next layer is where the work is. ## Portfolio ### Axiom Space - New Frontiers · Investor - Website: https://www.axiomspace.com Commercial space infrastructure company building the world's first private space station and developing commercial human spaceflight capabilities to enable a sustainable space economy. ### Blockdaemon - Infrastructure · Investor - Website: https://www.blockdaemon.com Institutional-grade blockchain infrastructure platform providing node management, staking, and API access across multiple blockchain networks. ### Caribou Biosciences - New Frontiers · Investor - Website: https://www.cariboubio.com CRISPR genome editing company developing allogeneic CAR-T and other cell therapies for oncology and autoimmune disease. Public on Nasdaq as CRBU. ### Chef - Infrastructure · Founder - Website: https://www.chef.io Open-source cloud infrastructure automation platform that became one of the foundational tools of the DevOps movement, used by thousands of organizations worldwide. ### CircleCI - Developer Platforms · Investor - Website: https://circleci.com Continuous integration and delivery platform that automates the build, test, and deploy process for software teams. ### Colimit - AI - Tools · Board Member - Website: https://colimit.ai AI platform that analyzes and fixes failed CI builds, extracting root causes from log noise and identifying relevant files through deep codebase analysis. ### Conjur - Security · Advisor - Website: https://www.conjur.org Secrets management and machine identity platform for DevOps teams, providing secure, policy-based access to secrets across cloud, container, and on-premise environments. ### Continue - AI - Tools · Board Member - Website: https://continue.dev Open-source AI coding assistant and continuous AI platform for the enterprise, giving development teams control over how AI integrates into their workflow. ### Eight Sleep - Consumer · Investor - Website: https://www.eightsleep.com Sleep technology company making smart mattresses and sleep systems that track biometrics, regulate temperature dynamically, and use AI to optimize sleep health for peak performance. ### Ellipsis Health - AI - Physical & Defense · Investor - Website: https://www.ellipsishealth.com AI platform that analyzes vocal biomarkers to screen for depression and anxiety, providing scalable, objective mental health assessment through routine clinical voice interactions. ### Epirus - AI - Physical & Defense · Investor - Website: https://www.epirusinc.com Defense technology company building high-power microwave systems, including Leonidas, to counter drones, swarms, and other electronic threats. ### Fastly - Infrastructure · Advisor - Website: https://www.fastly.com Edge cloud platform delivering fast, secure, and scalable digital experiences through a programmable content delivery network and edge computing. ### Figure AI - AI - Physical & Defense · Investor - Website: https://www.figure.ai AI robotics company building general-purpose humanoid robots for physical work across industrial, logistics, and home environments. ### Firefly - Consumer · Investor - Website: https://www.fireflyon.com On-vehicle digital advertising platform that equips rideshare and taxi drivers with smart screens for geofenced, real-time ad targeting while sharing revenue with drivers. ### Gradle - Developer Platforms · Investor - Website: https://gradle.com Build automation platform that accelerates developer productivity through advanced caching, build scans, and performance optimization for software builds. ### Groundcover - Data & Observability · Investor - Website: https://www.groundcover.com Cloud-native observability platform using eBPF to provide zero-overhead monitoring of Kubernetes workloads without code instrumentation. ### Honor - Consumer · Investor - Website: https://joinhonor.com Home care platform that connects trained professional caregivers with seniors needing in-home assistance, using technology and quality training to improve care outcomes and caregiver experience. ### Instacart - Consumer · Advisor - Website: https://www.instacart.com Online grocery delivery and pickup platform connecting consumers with personal shoppers and retail partners for same-day delivery. ### Joy - Consumer · Investor - Website: https://withjoy.com Wedding planning and registry platform that gives couples a single system for guest lists, websites, invitations, travel, and gifts from engagement through honeymoon. ### Kenn - AI - Tools · Advisor - Website: https://kenn.io Small team led by Wes McKinney (pandas, Apache Arrow, Ibis) building local-first developer tools for working with coding agents, including AgentsView, roborev, kata, and msgvault. ### Kentik - Data & Observability · Investor - Website: https://www.kentik.com Network observability platform that transforms network telemetry into real-time analytics, giving engineering and operations teams the ability to monitor, troubleshoot, and optimize network performance at scale. ### LaunchDarkly - Developer Platforms · Investor - Website: https://launchdarkly.com Feature management platform that lets engineering teams deploy code without releasing features, supporting controlled rollouts and experimentation at scale. ### LocalStack - Infrastructure · Investor - Website: https://localstack.cloud Cloud development platform that emulates AWS services locally, letting developers build and test cloud applications without deploying to the cloud. ### Memgraph - AI - Tools · Board Member - Website: https://memgraph.com Open-source graph database and early leader in GraphRAG, the essential data layer for AI applications that need relationship-rich context, real-time streaming, and millisecond-latency graph queries. ### Milestone - AI - Tools · Investor - Website: https://mstone.ai GenAI adoption and ROI platform that helps engineering teams measure the impact of AI tools on productivity, code quality, and development velocity. ### Mobot - AI - Physical & Defense · Board Member - Website: https://www.mobot.io AI-powered mobile app testing platform that uses robotic systems to test applications on real physical devices, ensuring quality across device and OS combinations. ### Musical AI - AI - Tools · Board Member - Website: https://www.wearemusical.ai Attribution and rights-aware AI infrastructure for media companies working with generative audio and music workflows. ### Netlify - Developer Platforms · Investor - Website: https://www.netlify.com Web development platform that pioneered the Jamstack architecture, letting developers build and deploy modern web applications with Git-based workflows. ### Omnitron Sensors - AI - Physical & Defense · Investor - Website: https://omnitronsensors.com Sensor technology company developing advanced sensing systems for industrial and defense applications requiring high precision and reliability in challenging environments. ### Orion Labs - AI - Physical & Defense · Founder - Website: https://www.orionlabs.io Voice-first communication platform for field teams and frontline workers, combining the Onyx wearable push-to-talk device with an intelligent agent platform that connects people, IoT devices, and AI across any distance. ### PagerDuty - Infrastructure · Advisor - Website: https://www.pagerduty.com Digital operations management platform that helps engineering teams detect, triage, and resolve incidents in real time. ### Particle - Developer Platforms · Board Member - Website: https://www.particle.io Integrated IoT Platform-as-a-Service combining developer hardware modules, cellular connectivity, and cloud management for building and deploying connected device products at scale. ### Prime Radiant - AI - Tools · Advisor - Website: https://www.primeradiant.com AI research lab creating Superpowers — an agent framework and development platform designed to accelerate the transition to agentic systems. ### Radar - Developer Platforms · Investor - Website: https://radar.com Developer-friendly, privacy-first geofencing and location data infrastructure for building location-aware experiences for enterprises and startups. ### Recce - AI - Tools · Board Member - Website: https://www.recce.io Data validation platform that helps data teams review and validate changes to data pipelines before they reach production. ### Replicated - Developer Platforms · Investor - Website: https://www.replicated.com Platform that lets software vendors deploy and manage their applications inside customer environments, supporting on-prem and private cloud distribution at scale. ### Sanity - Developer Platforms · Board Member - Website: https://www.sanity.io Composable content platform that treats content as structured data, letting teams build flexible content infrastructure for any digital experience. ### Shield AI - AI - Physical & Defense · Investor - Website: https://shield.ai AI company building autonomous pilots and intelligent systems for defense aircraft and platforms, allowing military systems to operate in GPS-denied and communication-denied environments. ### SignalWire - Developer Platforms · Investor - Website: https://signalwire.com Cloud communications platform built by the creators of FreeSWITCH, providing programmable video, voice, and messaging APIs for developers building real-time communication applications. ### SkySafe - AI - Physical & Defense · Investor - Website: https://www.skysafe.io Airspace security platform for drone detection, identification, tracking, and mitigation, built for defense, public safety, and critical infrastructure operators. ### Snyk - Security · Investor - Website: https://snyk.io Developer security platform that helps teams find and fix vulnerabilities in code, open source dependencies, containers, and infrastructure as code. ### Tailscale - Infrastructure · Investor - Website: https://tailscale.com Zero-config mesh VPN built on WireGuard that makes secure networking simple, replacing complex network infrastructure with identity-based connectivity. ### Tonic.ai - AI - Tools · Investor - Website: https://www.tonic.ai Synthetic data platform that generates realistic, de-identified datasets from production data for development, testing, and AI model training. ### TruSTAR - Security · Board Member - Website: https://www.trustar.co Threat intelligence platform that collected, normalized, and operationalized security intelligence across teams and tools to accelerate incident response. ### Unikraft - Infrastructure · Investor - Website: https://unikraft.io Open-source unikernel platform that builds ultra-lightweight, high-performance virtual machines tailored to run a single application with minimal overhead. ### Vibrant Labs - AI - Tools · Investor - Website: https://vibrantlabs.com Reinforcement learning environments and evaluation infrastructure for long-horizon AI agents, born from Ragas, the leading open-source LLM evaluation framework. ### Workbrew - Developer Platforms · Investor - Website: https://workbrew.com Enterprise Homebrew platform that gives IT and security teams centralized control over macOS package management while preserving developer productivity. ### Zymergen - New Frontiers · Investor - Website: https://zymergen.com Synthetic biology company that combined machine learning and automated lab systems to engineer microbes for producing novel materials and molecules. Acquired by Ginkgo Bioworks in 2022. ## Writing ### Announcing Greenroom: A private place to prepare your public commits - 2026-08-10 - Canonical: https://jesserobbins.com/blog/greenroom/ Writing your design and spec docs down generates rich context for you and your agents, and most of it does not belong in the public repo. Greenroom keeps it under git, in a private repo next to the public one. I have been using Jesse Vincent's incredible [Superpowers](https://github.com/obra/Superpowers) actively for months now, and it is what really unlocked agentic coding for me. I've built my own agentic harnesses, which I've been using to help me contribute to great open source projects like [agentsview](https://agentsview.io), [msgvault](https://msgvault.io), and [maclocal-api](https://github.com/scouzi1966/maclocal-api). Doing this well means writing and iterating on a lot of docs, tests, etc. While this is critical context for me and for my agents... I think everyone is still figuring out where all this new context belongs. At the same time, the contribution bar has risen dramatically. Open source maintainers expect great communication, small, clear commits, with finely refined-and-polished code. To me, checking all of my work, my design docs, test plans, and notes on my own thinking into a *public* repo feels... weirdly vulnerable and private. It reminds me of [stepping on to a stage to perform](https://jesserobbins.com/mentions/velocity-2012-changing-culture-force-awesome-oreilly/), and I only want to do that when I am ready. In the theatre and conferences, a [Green Room](https://en.wikipedia.org/wiki/Green_room) is the private space where performers prepare to take the stage. So I made **Greenroom: A private place to prepare your public commits** Greenroom creates and maintains two (or more) repos under a top level project folder and then manages them carefully. For my open source forks and contributions I have a repo-public (my fork), repo-private (my Greenroom directory). Sometimes I have several forks I'm managing, or pull in dependencies, etc. ## Install greenroom in claude code ```text /plugin marketplace add jesserobbins/greenroom /plugin install greenroom@jesserobbins ``` ## Or with skills.sh for everything else The standalone skill copies Greenroom straight into your project and works on Claude Code, Codex, Cursor, and the other 70-plus agents that [skills.sh](https://skills.sh/jesserobbins/greenroom) supports: ```text npx skills add jesserobbins/greenroom ``` ## Greenroom re-organizes your project in this layout ```text ~/src// # parent folder, not a git repo ├── AGENTS.md # orientation for any agent launched here ├── -public/ # public code repo (the thing on GitHub: the stage) └── -private/ # private notes repo (a separate private GitHub repo: the Greenroom) ├── AGENTS.md # private-side orientation ├── docs/ # design docs ├── notes/ # dated working notes ├── drafts/ # PR/issue/blog drafts ├── reviews/ # private notes on PRs └── research/ # transcripts, links, experiments ``` I expect most people will use Greenroom to contribute to other people's projects. This lets you build/iterate/refine/polish in private while keeping your notes until you are ready to push to public and issue that awesome PR. It is also great if you are the maintainer! I think in the end we will see maintainers want design docs and plans along with high quality PR submissions, and everyone is still figuring that out. Give it a try and let me know if it helps! [github.com/jesserobbins/greenroom](https://github.com/jesserobbins/greenroom). ## Lab ### What the Apple Neural Engine and Google's TPU tell us about the next decade of inference - Essay · 2026-06-08 - Canonical: https://jesserobbins.com/research/ane-tpu-deep-research/ A deep-research report on neural processors. What they are, how they differ from GPUs, and what Apple and Google have built across the Apple Neural Engine (ANE) and the TPU lineage.

A note from Jesse

I have been excited about NPUs since [Alex McNamara](https://www.linkedin.com/in/alexmcnamara/) first proposed using them for our on-device AI work at Orion Labs in 2019. Obviously I am a big believer in moving intelligence as close to the edge as possible. Many people are surprised to learn that there are supercomputer level Trillion Operation Per Second dedicated processors in most iPhones. I have returned to this hands-on recently in my contributions to [maclocal-api](https://github.com/scouzi1966/maclocal-api), an open-source project that puts Apple's on-device models behind an OpenAI-compatible HTTP API. Over the last several weeks I have added support for Apple's on-device [speech transcription and text-to-speech](https://github.com/scouzi1966/maclocal-api/pull/113), an [embeddings endpoint backed by Apple's NaturalLanguage framework](https://github.com/scouzi1966/maclocal-api/pull/119), and a set of [Vision endpoints for OCR, barcodes, image classification, and saliency](https://github.com/scouzi1966/maclocal-api/pull/114). When I talk to people about the research and commits I have made, I keep having to explain the Apple Neural Engine (ANE) and TPUs each time. I realized I should just run a Claude deep-research pass and share the output.
## What a neural processor actually is
Illustrative side-by-side comparison of how a CPU, a GPU, and an NPU allocate their on-chip area, and what each is best at. The CPU dedicates roughly half its area to control and scheduling logic, with caches and a small number of wide ALUs, and is best at branchy logic, control flow, and the operating system. The GPU dedicates most of its area to thousands of small parallel cores and is best at graphics, simulation, and large-batch model training. The NPU dedicates the majority of its area to a dense multiply-accumulate array and on-chip SRAM, and is best at low-power, always-on, on-device neural network inference.
How a CPU, a GPU, and an NPU spend their silicon, and what each is best at. Schematic and illustrative; real chips vary widely. The point is the design intent.
A neural processor (commonly called an NPU, for neural processing unit) is a piece of silicon designed for one specific job: running the math at the heart of a neural network. That math is overwhelmingly multiply-accumulate operations on small matrices and tensors of fixed-point or low-precision floating-point numbers. An NPU is built around that one job and very little else.
A visual explanation of the matrix multiplication operation that dominates neural network inference. One row of input matrix A is multiplied element-wise with one column of input matrix B, and the products are summed to produce a single cell of the output matrix C. The example shows the row 2, 3, 1, 4 multiplied by the column 5, 2, 7, 1, producing 27. A neural network does billions of these operations per inference pass; an NPU is silicon built to do them in parallel with as little overhead as possible.
What the math actually looks like. One row of A, one column of B, multiplied element-wise and summed into one cell of C. That is a single multiply-accumulate. An NPU runs thousands of these in parallel.
The clearest way to see what an NPU is is to compare it to the three things it sits next to in a modern system. **CPU.** A general-purpose processor. It runs arbitrary control flow, branchy logic, the operating system, and everything else. It is the most flexible compute in the system and the least efficient for tensor math, because almost none of its transistors are dedicated to wide parallel multiplies. **GPU.** A massively parallel processor originally built for graphics, retrofitted very successfully for machine learning. GPUs are good at the same wide parallel arithmetic that neural networks need, which is why they powered the deep-learning era. They are programmable through general-purpose shading languages (Metal, CUDA, Vulkan compute) and can run almost anything that maps to thousands of parallel threads. Training large models on GPUs at scale is what built modern AI. **NPU.** A fixed-function accelerator built for one shape of computation: dense low-precision matrix multiplies and convolutions, typically with built-in support for the activations, normalizations, and quantization steps that neural networks need. NPUs are not designed to be fully programmable. They are designed to do the inner loop of an inference pass at very low power and very high throughput. **The trade-off** is straightforward. A GPU can run a neural network and a fluid simulation and a fragment shader. An NPU can run a neural network. In exchange for giving up that generality, an NPU spends nearly all of its transistor budget on the multiplier arrays, the on-chip SRAM that feeds them, and the data-movement plumbing between the two. For the workloads it is built for, the result is more operations per second per watt than a GPU of comparable area. That last phrase, *per watt*, is where NPUs matter most. A datacenter GPU can draw 400 to 700 watts; a phone or a watch cannot. If you want a model to run continuously on a battery-powered device, listening for a wake word or watching a camera feed, the power envelope is measured in milliwatts, not watts. That budget is what makes a fixed-function accelerator the right tool: you give up programmability to spend the silicon on doing the one job efficiently. Both of the chips this report is about, the Apple Neural Engine and Google's TPU, are NPUs in this sense. Both are built around large arrays of multiply-accumulate units. Both are paired with software runtimes that handle the data movement, quantization, and model compilation that turn a high-level model description into something the silicon can execute. They sit at different points on the size and power curve, one designed for the inside of a phone, the other for the inside of a datacenter rack, but the underlying design philosophy is the same: dedicate the silicon to the math the network actually does. For a longer, more visual treatment of how systolic arrays and dedicated tensor silicon work, the videos in the "Watch" section below are good starting points. ## A. What the Apple Neural Engine actually is The ANE has shipped in every Apple-designed system-on-chip since the A11 Bionic in September 2017, the chip Apple introduced alongside the iPhone X. At launch Apple described the original ANE as a dual-core block capable of 600 billion operations per second, used initially to power Face ID, Animoji, and real-time image processing ([CNBC, Sep 12 2017](https://www.cnbc.com/2017/09/12/apple-unveils-a11-bionic-neural-engine-ai-chip-in-iphone-x.html)). It is a fixed-function NPU, not a general-purpose accelerator. It is optimized for FP16 convolutional and matrix-multiplication workloads, the operations that dominate vision models and modern transformer inference. **Throughput growth, 2017 to 2021.** Apple's published peak figures show a roughly 26-fold jump in four years: the A11's 0.6 TFLOPS gave way to a 16-core ANE in the A15 Bionic (2021) that Apple rates at 15.8 TFLOPS. Treat these as vendor peak numbers against vendor reference benchmarks; they are accurate as published but should be framed that way. **The distilbert case study.** In 2022, Apple's machine learning research team published a reference implementation showing how to adapt a Hugging Face transformer model to run efficiently on the ANE. With the optimizations applied, Apple reported the model running up to 10× faster and using 14× less peak memory, with an end-to-end latency of 3.47 ms at 0.454 W on an iPhone 13 (sequence length 128, batch size 1) ([Deploying Transformers on the Apple Neural Engine, Apple ML Research, 2022](https://machinelearning.apple.com/research/neural-engine-transformers); reference code: [apple/ml-ane-transformers](https://github.com/apple/ml-ane-transformers); optimized model weights: [apple/ane-distilbert-base-uncased-finetuned-sst-2-english on Hugging Face](https://huggingface.co/apple/ane-distilbert-base-uncased-finetuned-sst-2-english)). **The software stack predates the LLM wave.** Most of the framing around on-device AI treats the current moment as new. Apple's developer surface for it is not. The relevant milestones, in order: - **Core ML (2017).** Apple's high-level inference framework, the public path to running ML models on Apple Silicon. Documentation at [developer.apple.com/documentation/coreml](https://developer.apple.com/documentation/coreml). - **Natural Language framework (iOS 12, WWDC 2018).** On-device language identification, tokenization, lemmatization, part-of-speech tagging, and named-entity recognition ([Introducing Natural Language Framework, WWDC 2018, Session 713](https://nonstrict.eu/wwdcindex/wwdc2018/713/)). - **WWDC 2019 additions.** On-device sentiment analysis across seven languages and word embeddings, used together for in-app search and similarity ([Advances in Natural Language Framework, WWDC 2019, Session 232](https://developer.apple.com/videos/play/wwdc2019/232/); session transcript at [ASCIIwwdc](https://asciiwwdc.com/2019/sessions/232)). Apple's documentation describes the sentiment classifier as "hardware activated" on supported devices; that is Apple's phrase and it does not explicitly name the ANE, so do not upgrade it to a guaranteed-ANE claim. - **MLX (2023).** Apple's array framework for machine learning on Apple Silicon, with a unified memory model and a NumPy-like Python API. This is the path for training and fine-tuning workflows on Apple hardware, distinct from Core ML's inference focus ([ml-explore/mlx](https://github.com/ml-explore/mlx)). - **Foundation Models framework (iOS 26, WWDC 2025).** A Swift API for tapping Apple Intelligence's on-device ~3B-parameter model in roughly three lines of code, with guided generation, streaming, and tool calls ([Foundation Models, Apple Developer Documentation](https://developer.apple.com/documentation/FoundationModels); [Meet the Foundation Models framework, WWDC 2025, Session 286](https://developer.apple.com/videos/play/wwdc2025/286/); [Deep dive into the Foundation Models framework, WWDC 2025, Session 301](https://developer.apple.com/videos/play/wwdc2025/301/)).
Apple's introduction to the Foundation Models framework in iOS 26 (WWDC 2025, Session 286). The three-line "hello world" for on-device LLM use lives at around the four-minute mark. Mirror on the official Apple YouTube channel.
Nine years of public, mature on-device ML infrastructure. The headlines moved on; the stack kept growing.
For a current end-to-end orientation across Core ML, MLX, Create ML, and the Vision, Natural Language, and Speech frameworks, Apple's WWDC 2024 overview is the single best place to start. Mirror on the official Apple YouTube channel.
## B. How Apple has chosen to expose it Apple's developer surface for the ANE is shaped by a clear design choice: the operating system, not the application, decides where neural work runs. The result is a runtime that lets app developers ship machine learning features without having to reason about which compute unit to target. **Core ML is the public path, and it is a scheduler.** At runtime, Core ML inspects a model and decides, layer by layer, whether each one runs best on the CPU, the GPU, or the ANE. The developer ships the model; Apple's runtime handles the placement. As Apple silicon evolves, the same model gets faster without the app having to ship an update, because the scheduler knows what the new hardware can do. The configuration surface reflects that philosophy. `MLComputeUnits` exposes four preferences: `all`, `cpuOnly`, `cpuAndGPU`, and `cpuAndNeuralEngine` ([MLComputeUnits, Apple Developer Documentation](https://developer.apple.com/documentation/coreml/mlcomputeunits)). The defaults are tuned for the common case. The narrower options exist for the cases where a developer has a reason to constrain the runtime, for example to keep the GPU free for graphics work or to validate behavior under a specific compute path ([MLComputeUnits.cpuAndNeuralEngine, Apple Developer Documentation](https://developer.apple.com/documentation/coreml/mlcomputeunits/cpuandneuralengine)). The most concrete public illustration is Apple's own WWDC 2022 session, "Optimize your Core ML usage." Presenters run a YOLOv3 object-detection model under the `all` setting and use Xcode's performance report to inspect, layer by layer, where Core ML placed the work. The result for that model: 54 layers on the GPU and 32 on the ANE. The developer set a single preference; the runtime made 86 informed placement decisions on their behalf, including the ones that fit the ANE's matrix-multiply units best ([Optimize your Core ML usage, WWDC 2022, Session 10027](https://developer.apple.com/videos/play/wwdc2022/10027/)).
Apple's "Optimize your Core ML usage" from WWDC 2022 (Session 10027). The YOLOv3 layer-distribution walkthrough is the clearest existing demonstration of how Core ML decides where each layer of a model runs. Mirror on the official Apple YouTube channel.
**Training has its own path.** Core ML and the ANE focus on inference. For training and fine-tuning on Apple Silicon, Apple ships MLX, a NumPy-style array framework with a unified memory model that targets the CPU and GPU directly, and PyTorch on Apple Silicon runs through Metal Performance Shaders. The split is deliberate: the ANE is built for the inference inner loop, and the GPU is the natural home for backprop. Apple's WWDC 2025 sessions on MLX go deep on how to use it for LLM fine-tuning on consumer Macs ([Get started with MLX for Apple silicon, WWDC 2025, Session 315](https://developer.apple.com/videos/play/wwdc2025/315/); [Explore large language models on Apple silicon with MLX, Session 298](https://developer.apple.com/videos/play/wwdc2025/298/)). **A growing community is exploring the ANE in the open.** While Apple's published material on the ANE itself is limited and the official channels for working with it are intentionally high-level, an active community of independent researchers and Apple itself have been building tools, documentation, benchmarks, and reference implementations around the chip. The most useful entry points to that work: - **[hollance/neural-engine](https://github.com/hollance/neural-engine)** — Matthijs Hollemans's comprehensive community documentation of ANE behavior, performance characteristics, and supported operations. The single best existing resource on the ANE. - **[mdaiter/ane](https://github.com/mdaiter/ane)** — early reverse engineering with working Python and Objective-C samples, documenting the ANECompiler framework and IOKit dispatch. - **[eiln/ane](https://github.com/eiln/ane)** — a reverse-engineered Linux driver for the ANE from the Asahi Linux project, providing insight into the kernel-level interface. - **[apple/ml-ane-transformers](https://github.com/apple/ml-ane-transformers)** — Apple's own reference implementation of transformers optimized for the ANE, confirming design patterns like channel-first layout and a preference for 1×1 convolutions. - **[Anemll/Anemll](https://github.com/Anemll/Anemll)** (pronounced "animal," for "Artificial Neural Engine Machine Learning Library") — an open-source project focused on running large language models directly on the ANE, with a single-file conversion-and-inference pipeline for LLaMA, Qwen, Qwen 2.5, and Gemma 3 architectures, plus a companion benchmarking tool at [Anemll/anemll-bench](https://github.com/Anemll/anemll-bench). - **[maderix/ANE](https://github.com/maderix/ANE)** — research into training on the M4 ANE, building a custom compute graph with a backward pass by talking to the lower-level frameworks inside `AppleNeuralEngine.framework`. The author is careful about the caveats: the proof of concept runs at roughly 5–9% of peak ANE utilization, and the methodology depends on undocumented APIs ([Inside the M4 Apple Neural Engine, Part 1: Reverse Engineering](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine)). The natural reading is that Apple is choosing the rate at which to open up programmability, and the hardware will let them go further when they decide to. In the meantime the community has done a lot of careful work to map what is already there. The overall shape: Apple has built a mature on-device ML stack where the runtime handles the hard scheduling problem so applications do not have to. That choice has shipped a lot of working machine learning into a lot of pockets. ## C. The TPU lineage and why it matters Step back from Apple for a moment and the broader picture is that dedicated tensor silicon has been in production at industrial scale for over a decade, and the architectural ideas have been migrating from datacenters to the edge. **The first-generation TPU.** Google's first TPU went into production inside Google datacenters in 2015. Norman Jouppi and colleagues described it in detail at ISCA 2017 in a paper that has become the standard reference: a custom ASIC built around a 256×256 systolic-array matrix-multiply unit, with 65,536 8-bit MACs, a peak throughput of 92 TeraOps/second, and 28 MiB of software-managed on-chip memory ([In-Datacenter Performance Analysis of a Tensor Processing Unit, Jouppi et al., ISCA 2017, arXiv:1704.04760](https://arxiv.org/abs/1704.04760); [ACM proceedings entry](https://dl.acm.org/doi/10.1145/3079856.3080246)). **Performance, with caveats.** The same paper reports the first-generation TPU as 15× to 30× faster than a contemporary Intel Haswell CPU and a contemporary NVIDIA K80 GPU on Google's production inference workloads, with 30× to 80× better performance-per-watt. These are Google-published numbers against Google's reference workloads; they were peer-reviewed and have aged well, but they are vendor benchmarks and should be framed that way. For a retrospective tour across four TPU generations from one of the paper's co-authors, see David Patterson's [*A Decade of Machine Learning Accelerators: Lessons Learned and Carbon Footprint*](https://www.cs.ucla.edu/wp-content/uploads/cs/PATTERSON-10-Lessons-4-TPU-gens-CO2e-45-minutes.pdf) (slide deck, 2022), and the authors' own [*Ten Lessons from Three Generations Shaped Google's TPUv4i*](https://dl.acm.org/doi/abs/10.1109/ISCA52012.2021.00010) (ISCA 2021). **The lineage extends to the edge.** Google has carried the tensor-accelerator design philosophy down into wearable- and hearable-class hardware. Coral NPU, announced in 2025, is an open-source 32-bit RISC-V design with a vector co-processor implementing the RVV v1.0 vector ISA and a quantized outer-product MAC engine that processes 8-bit operations into 32-bit results. It is targeted at ultra-low-power, always-on edge AI, including smart watches and AR glasses ([Coral NPU datasheet, Google for Developers](https://developers.google.com/coral/guides/hardware/datasheet); [Introducing Coral NPU, Google Developers Blog](https://developers.googleblog.com/en/introducing-coral-npu-a-full-stack-platform-for-edge-ai/); source: [google-coral/coralnpu on GitHub](https://github.com/google-coral/coralnpu)).
Google's official Coral NPU architecture diagram. A scalar RISC-V core sits at the top, dispatching to a vector execution unit (implementing RVV v1.0) and a matrix execution unit (the outer-product multiply-accumulate engine). All three share access to tightly coupled instruction and data memory and an external memory interface.
Coral NPU architecture: a scalar RISC-V core dispatches to a vector execution unit and a matrix (MAC) execution unit, sharing a tightly coupled memory hierarchy. Source: Coral NPU Architecture overview, Google for Developers (Apache 2.0).
A 2015 datacenter ASIC and a 2025 open-source RISC-V NPU for hearables are the same idea a decade apart: matrix-multiply silicon got faster, smaller, more efficient, and ended up everywhere. ## D. What this tells us about the next decade of inference **Inference is migrating outward from the datacenter.** Google's first-generation TPU went into production inside Google datacenters in 2015 ([Jouppi et al., ISCA 2017, arXiv:1704.04760](https://arxiv.org/abs/1704.04760)). Ten years later, the same company has published a 32-bit RISC-V NPU targeted at smart watches and AR glasses ([Introducing Coral NPU, Google Developers Blog](https://developers.googleblog.com/en/introducing-coral-npu-a-full-stack-platform-for-edge-ai/); source at [google-coral/coralnpu](https://github.com/google-coral/coralnpu)). Apple has shipped an NPU in every iPhone since the A11 Bionic in September 2017 ([CNBC, Sep 12 2017](https://www.cnbc.com/2017/09/12/apple-unveils-a11-bionic-neural-engine-ai-chip-in-iphone-x.html)), and at WWDC 2025 made a roughly 3-billion-parameter on-device model available to any third-party developer through the Foundation Models framework ([Foundation Models, Apple Developer Documentation](https://developer.apple.com/documentation/FoundationModels); [Meet the Foundation Models framework, WWDC 2025, Session 286](https://developer.apple.com/videos/play/wwdc2025/286/)). Qualcomm describes its current mobile NPUs as "designed from the ground up for accelerating AI inference at low power" and ships them across the Snapdragon line ([Hexagon NPU, Qualcomm](https://www.qualcomm.com/products/technology/processors/hexagon); [Snapdragon for AI on-device, Qualcomm](https://www.qualcomm.com/products/technology/artificial-intelligence)). Across three independent vendors, the public record points in the same direction: a growing share of inference work runs on dedicated silicon outside the datacenter. **Most of the headroom available on chips already in pockets is in software.** Apple's developer-facing path to the ANE has widened in steady increments since 2017: Core ML in 2017 ([Core ML, Apple Developer Documentation](https://developer.apple.com/documentation/coreml)), the Natural Language framework in 2018 ([WWDC 2018, Session 713](https://nonstrict.eu/wwdcindex/wwdc2018/713/)), MLX in 2023 ([ml-explore/mlx](https://github.com/ml-explore/mlx)), the Foundation Models framework in 2025 ([Deep dive into the Foundation Models framework, WWDC 2025, Session 301](https://developer.apple.com/videos/play/wwdc2025/301/)), and updated MLX guidance for LLM work on Apple silicon at the same WWDC ([Get started with MLX for Apple silicon, WWDC 2025, Session 315](https://developer.apple.com/videos/play/wwdc2025/315/); [Explore large language models on Apple silicon with MLX, WWDC 2025, Session 298](https://developer.apple.com/videos/play/wwdc2025/298/)). Independent research suggests the silicon can do more than the current public path expresses: the maderix proof of concept runs at 5–9% of peak ANE utilization through reverse-engineered private APIs ([Inside the M4 Apple Neural Engine, Part 1, maderix](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine); benchmarks in [Part 2](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615); code at [maderix/ANE on GitHub](https://github.com/maderix/ANE)). The gap between what the chip can do and what the developer-facing path can reach is large enough that closing it is itself a source of performance. The same pattern shows up on the datacenter side: Patterson's retrospective across four TPU generations attributes much of the improvement to compiler and software-stack work rather than process shrinks alone ([*A Decade of Machine Learning Accelerators*, David Patterson, 2022](https://www.cs.ucla.edu/wp-content/uploads/cs/PATTERSON-10-Lessons-4-TPU-gens-CO2e-45-minutes.pdf); [*Ten Lessons from Three Generations Shaped Google's TPUv4i*, Jouppi et al., ISCA 2021](https://dl.acm.org/doi/abs/10.1109/ISCA52012.2021.00010)). **The fixed-function trade has held up under scrutiny.** Google's published numbers for the first-generation TPU, 15× to 30× faster than its CPU and GPU contemporaries at 30× to 80× better performance-per-watt on Google production inference workloads, were peer-reviewed in 2017 and have been revisited in the literature several times since ([Jouppi et al., ISCA 2017](https://arxiv.org/abs/1704.04760); [retrospective, Jouppi 2023](https://bpb-us-w2.wpmucdn.com/sites.coecis.cornell.edu/dist/7/587/files/2023/06/Jouppi_2017_In_Datacenter.pdf)). Apple's distilbert reference implementation reports 3.47 ms end-to-end latency at 0.454 W on an iPhone 13 with the ANE-tuned model ([Deploying Transformers on the Apple Neural Engine, Apple ML Research, 2022](https://machinelearning.apple.com/research/neural-engine-transformers); code at [apple/ml-ane-transformers](https://github.com/apple/ml-ane-transformers); model weights at [apple/ane-distilbert-base-uncased-finetuned-sst-2-english](https://huggingface.co/apple/ane-distilbert-base-uncased-finetuned-sst-2-english)). The shape of workload these chips are built for, dense low-precision matrix multiplies and convolutions, has been stable across the ten-year window between those two results. **The application-facing API for on-device models is converging on something close to a library call.** Apple's Foundation Models framework reduces invoking a 3B-parameter model from Swift to roughly three lines of code, with the "hello world" at around the four-minute mark of the WWDC 2025 introduction ([WWDC 2025, Session 286](https://developer.apple.com/videos/play/wwdc2025/286/)). Google's mobile path takes the same shape: Gemini Nano runs through Android's AICore service, and applications reach it through ML Kit's GenAI APIs and the Google AI Edge SDK ([Gemini Nano via AICore, Android Developers](https://developer.android.com/ai/aicore); [ML Kit GenAI APIs, Google for Developers](https://developers.google.com/ml-kit/genai); [Google AI Edge SDK overview](https://ai.google.dev/edge)). The cross-platform on-device path through ONNX Runtime exposes a single inference API that dispatches to whichever NPU or GPU is present on the host ([ONNX Runtime documentation](https://onnxruntime.ai/docs/); [Hugging Face transformers.js for in-browser inference](https://huggingface.co/docs/transformers.js)). The 2017-to-2025 arc across Core ML, MLX, AICore, and ONNX Runtime is the abstraction layer rising in roughly the same direction across vendors. **Open access to the silicon is widening on both vendor and community sides.** Google has published Coral NPU as an open-source 32-bit RISC-V design with a vector co-processor implementing the [RVV v1.0](https://github.com/riscv/riscv-v-spec) vector ISA and a quantized outer-product MAC engine ([Coral NPU datasheet, Google for Developers](https://developers.google.com/coral/guides/hardware/datasheet); source at [google-coral/coralnpu](https://github.com/google-coral/coralnpu)). On the Apple side, several independent research projects have produced public tooling for studying and exercising the ANE, including [hollance/neural-engine](https://github.com/hollance/neural-engine), [Anemll/Anemll](https://github.com/Anemll/Anemll) with its [companion benchmark suite](https://github.com/Anemll/anemll-bench), [maderix/ANE](https://github.com/maderix/ANE), [mdaiter/ane](https://github.com/mdaiter/ane), and the Asahi Linux project's reverse-engineered Linux driver at [eiln/ane](https://github.com/eiln/ane). On the Apple side, the published interface remains Core ML's coarse-grained scheduler ([MLComputeUnits, Apple Developer Documentation](https://developer.apple.com/documentation/coreml/mlcomputeunits)); the community work above is what has made finer-grained study of the chip possible in public. ## Summary of the hardest numbers The three figures worth carrying out of this report, all from primary vendor or peer-reviewed sources: 1. **ANE throughput grew roughly 26× in four years.** From 0.6 TFLOPS in the A11 Bionic (2017) to 15.8 TFLOPS in the 16-core ANE of the A15 Bionic (2021). 2. **Apple's distilbert reference implementation runs up to 10× faster and uses 14× less peak memory after ANE optimizations**, with 3.47 ms latency at 0.454 W on an iPhone 13 ([Apple ML Research, 2022](https://machinelearning.apple.com/research/neural-engine-transformers)). 3. **The first-generation TPU was 15× to 30× faster and 30× to 80× more performance-per-watt than its CPU and GPU contemporaries** on Google's production inference workloads ([Jouppi et al., ISCA 2017](https://arxiv.org/abs/1704.04760)). ## Caveats - Apple and Google performance numbers in this post are vendor benchmarks against vendor-chosen reference workloads. They are accurate as published, and where they have been independently scrutinized (the TPU paper most of all) they have held up. Treat them as the floor of what the architecture can do under conditions the vendor chose, not as a guarantee for arbitrary models. - Apple's "hardware activated" phrasing around on-device sentiment analysis is Apple's own. It implies hardware acceleration on supported devices and does not explicitly name the ANE. Reporting it without that nuance would be an upgrade of the claim Apple actually made. - The argument that the constraint on ANE adoption is software rather than hardware comes from the maderix project. The proof of concept supporting it runs at 5–9% of peak ANE utilization and depends on reverse-engineered private APIs. The hardware-vs-software framing is a reasonable inference from that work; it should be attributed to its source, not asserted as a settled finding. ## Sources Primary sources, in order of first appearance: - [Apple unveils A11 Bionic neural engine AI chip in iPhone X — CNBC, Sep 12 2017](https://www.cnbc.com/2017/09/12/apple-unveils-a11-bionic-neural-engine-ai-chip-in-iphone-x.html) - [Deploying Transformers on the Apple Neural Engine — Apple Machine Learning Research, 2022](https://machinelearning.apple.com/research/neural-engine-transformers) - [apple/ml-ane-transformers — reference implementation on GitHub](https://github.com/apple/ml-ane-transformers) - [apple/ane-distilbert-base-uncased-finetuned-sst-2-english — optimized model on Hugging Face](https://huggingface.co/apple/ane-distilbert-base-uncased-finetuned-sst-2-english) - [Core ML — Apple Developer Documentation](https://developer.apple.com/documentation/coreml) - [Introducing Natural Language Framework — WWDC 2018, Session 713](https://nonstrict.eu/wwdcindex/wwdc2018/713/) - [Advances in Natural Language Framework — WWDC 2019, Session 232](https://developer.apple.com/videos/play/wwdc2019/232/) - [Advances in Natural Language Framework — ASCIIwwdc transcript](https://asciiwwdc.com/2019/sessions/232) - [ml-explore/mlx — MLX framework on GitHub](https://github.com/ml-explore/mlx) - [Foundation Models — Apple Developer Documentation](https://developer.apple.com/documentation/FoundationModels) - [Meet the Foundation Models framework — WWDC 2025, Session 286](https://developer.apple.com/videos/play/wwdc2025/286/) - [Deep dive into the Foundation Models framework — WWDC 2025, Session 301](https://developer.apple.com/videos/play/wwdc2025/301/) - [MLComputeUnits — Apple Developer Documentation](https://developer.apple.com/documentation/coreml/mlcomputeunits) - [MLComputeUnits.cpuAndNeuralEngine — Apple Developer Documentation](https://developer.apple.com/documentation/coreml/mlcomputeunits/cpuandneuralengine) - [Optimize your Core ML usage — WWDC 2022, Session 10027](https://developer.apple.com/videos/play/wwdc2022/10027/) - [Inside the M4 Apple Neural Engine, Part 1: Reverse Engineering — maderix](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine) - [maderix/ANE — training on the ANE via private APIs, on GitHub](https://github.com/maderix/ANE) - [Anemll/Anemll — ANEMLL, open-source LLM inference on the Apple Neural Engine](https://github.com/Anemll/Anemll) - [Anemll/anemll-bench — companion benchmarking suite for the ANE](https://github.com/Anemll/anemll-bench) - [hollance/neural-engine — Matthijs Hollemans, "Everything we actually know about the Apple Neural Engine"](https://github.com/hollance/neural-engine) - [mdaiter/ane — early ANE reverse engineering with Python and Objective-C samples](https://github.com/mdaiter/ane) - [eiln/ane — reverse-engineered Linux driver for the ANE, Asahi Linux project](https://github.com/eiln/ane) - [In-Datacenter Performance Analysis of a Tensor Processing Unit — Jouppi et al., ISCA 2017, arXiv:1704.04760](https://arxiv.org/abs/1704.04760) - [In-Datacenter Performance Analysis of a Tensor Processing Unit — ACM Digital Library](https://dl.acm.org/doi/10.1145/3079856.3080246) - [Coral NPU datasheet — Google for Developers](https://developers.google.com/coral/guides/hardware/datasheet) - [Introducing Coral NPU: A full-stack platform for Edge AI — Google Developers Blog](https://developers.googleblog.com/en/introducing-coral-npu-a-full-stack-platform-for-edge-ai/) - [google-coral/coralnpu — Coral NPU source on GitHub](https://github.com/google-coral/coralnpu) - [Hexagon NPU — Qualcomm](https://www.qualcomm.com/products/technology/processors/hexagon) - [Snapdragon for on-device AI — Qualcomm](https://www.qualcomm.com/products/technology/artificial-intelligence) - [Gemini Nano via Android AICore — Google for Developers](https://developer.android.com/ai/aicore) - [ML Kit GenAI APIs — Google for Developers](https://developers.google.com/ml-kit/genai) - [Google AI Edge SDK overview — Google AI for Developers](https://ai.google.dev/edge) - [ONNX Runtime documentation](https://onnxruntime.ai/docs/) - [Hugging Face transformers.js — in-browser inference documentation](https://huggingface.co/docs/transformers.js) - [RISC-V Vector Extension (RVV) specification — riscv/riscv-v-spec on GitHub](https://github.com/riscv/riscv-v-spec) Videos and supplementary materials: - [Optimize your Core ML usage — WWDC 2022, Session 10027, on Apple Developer](https://developer.apple.com/videos/play/wwdc2022/10027/) · [YouTube mirror](https://www.youtube.com/watch?v=THXq071qZ6E) - [Meet the Foundation Models framework — WWDC 2025, Session 286, on Apple Developer](https://developer.apple.com/videos/play/wwdc2025/286/) · [YouTube mirror](https://www.youtube.com/watch?v=mJMvFyBvZEk) - [Explore machine learning on Apple platforms — WWDC 2024, on Apple Developer](https://developer.apple.com/videos/play/wwdc2024/10223/) · [YouTube mirror](https://www.youtube.com/watch?v=p_hyo2FRil4) - [Get started with MLX for Apple silicon — WWDC 2025, Session 315](https://developer.apple.com/videos/play/wwdc2025/315/) - [Explore large language models on Apple silicon with MLX — WWDC 2025, Session 298](https://developer.apple.com/videos/play/wwdc2025/298/) - [A Decade of Machine Learning Accelerators — David Patterson, slide deck, 2022](https://www.cs.ucla.edu/wp-content/uploads/cs/PATTERSON-10-Lessons-4-TPU-gens-CO2e-45-minutes.pdf) - [Ten Lessons from Three Generations Shaped Google's TPUv4i — Jouppi et al., ISCA 2021](https://dl.acm.org/doi/abs/10.1109/ISCA52012.2021.00010) - [Retrospective on the original TPU paper — Jouppi, 2023 PDF](https://bpb-us-w2.wpmucdn.com/sites.coecis.cornell.edu/dist/7/587/files/2023/06/Jouppi_2017_In_Datacenter.pdf) - [What the Hell is a Neural Engine? — Greg Gant, 2024](https://blog.greggant.com/posts/2024/06/24/what-the-hell-is-an-apple-neural-engine.html) - [Inside the M4 Apple Neural Engine, Part 2: Benchmarks — maderix](https://maderix.substack.com/p/inside-the-m4-apple-neural-engine-615) ### How to run vector search on a website in the browser with no server (like I do) - Code story · 2026-06-06 - Canonical: https://jesserobbins.com/research/site-search-hybrid-in-browser/ Here's how to do semantic search on your website without a server. Open the search box and the page downloads a small embedding model and a 160 KB vector index; from then on every query is answered locally in the browser, fused with keyword search. What I built, what to install. ### Good Fences make Good Agents: sandvault + /sv skill (part 1 of n) - Code story · 2026-05-24 - Canonical: https://jesserobbins.com/research/sandvault-simple-sandbox-for-agents/ A simple solution to working with agents that cannot be trusted to run 'ls'. Install with brew. Works in seconds. I added a few features. ## Mentions ### The Seed 100: The Best Early-Stage Investors of 2026 - Business Insider · Article · 2026-05-18 - Original: https://www.businessinsider.com/seed-100-best-early-stage-vc-investors-2026-5#2.-jesse-robbins - Canonical: https://jesserobbins.com/mentions/seed-100-businessinsider-2026/ > "I look for founders who want to build the operating system for entire industries. This requires extraordinary taste, grit, drive, and a vision for the future." > — Jesse Robbins Business Insider named Jesse Robbins to its 2026 Seed 100, its annual ranking of the best early-stage venture investors. Robbins invests in AI developer tools and infrastructure, with recent investments including Figure AI, Shield AI, Instacart, and Fastly. He cofounded Chef and the DevOps movement. ### Sunil Dhaliwal, Founder of Amplify Partners, on Lessons from 27 Years in VC - We The Builders · Podcast · 2026-05-21 - Original: https://podcasts.apple.com/us/podcast/e30-sunil-dhaliwal-founder-of-amplify-partners-on/id1829681453?i=1000768983982 - Canonical: https://jesserobbins.com/mentions/sunil-dhaliwal-amplify-we-the-builders/ > "Fastly was an introduction through a former CEO of mine, Jesse Robbins … Jesse said whenever Artur Bergman quits his job, he's going to start something. I would back that guy … Artur is fucking brilliant. Believe what he says … And that is probably the single biggest success in the history of the firm." > — Sunil Dhaliwal, on We The Builders My friend Sunil Dhaliwal has proven to be one of the most successful investors in our space. Two far-ranging hours on We The Builders, from the dot-com bubble to where the AI boom is really constrained to building Amplify. Go listen to the whole thing. This is an extraordinary two-hour conversation with my friend Sunil Dhaliwal, who has proven to be one of the most successful investors in our space. He goes far and wide here, from the dot-com bubble to the AI boom, early-stage firm investment dynamics, and why he built Amplify in the first place. It is a really great interview, and worth a listen or a watch. Along the way he tells some stories about a few introductions I made for him during a particularly exciting time in infrastructure. Sunil led my Series B at Chef while I was CEO and also still Chair of the Velocity Conference. Sunil made it clear he wanted to run (and pay for) the bar at my speaker parties, and expected me to introduce him to the best people. That night I introduced him to my friends Artur Bergman and Simon Wistow who were just founding Fastly. This, of course, turned out to be one of Sunil's (and my) best investments. (Sunil also met the founders of Datadog at Velocity the same way.) Definitely a trick I learned to use myself over the years. One painful moment from this part of the interview was Suffiyan being surprised that "O'Reilly did conferences". Yes, they did... and we all miss them. Long and excellent conversation for anyone who is interested in venture. Worth every minute. Go listen. ### Call a savepoint - LinkedIn · LinkedIn · 2026-05-07 - Original: https://www.linkedin.com/posts/jesserobbins_hey-feeling-depleted-after-working-with-activity-7458153361582751744-9PbM - Canonical: https://jesserobbins.com/mentions/call-a-savepoint-ai-coding-compulsive-engagement-linkedin/ > "The pull you feel toward 'just one more iteration' is not evidence that the work needs more time. It is evidence that the schedule of small wins has trained your brain to expect another one." > — Jesse Robbins Working with AI engages the same dopamine machinery as slot machines. The hollow feeling at the end of a 2.3B-token week is the loop doing what loops like this do. The fix is a savepoint. > Hey... feeling depleted after working with AI? > > It is not a flaw in you. It is the predictable result of two things: this work has the reward loops of a game and the cognitive cost of supervising it is much higher than it looks. Together they are an addictive loop to keep you in the seat past the point where you should have walked away wondering why you feel hollowed out by a day that produced so much output. > > I learned this after hitting 2.3B tokens for the week & #11 on the tokenmaxxing leaderboard. Everyone doing extraordinary things with AI right now, the tokenmaxxers, running multiple billions of tokens a month across Claude, Codex, and Gemini... they all know this. They are the ones building the most ambitious systems, shipping the most code, producing the most remarkable output. > > Working with AI engages the same dopamine machinery as slot machines, video games, and social feeds. You write a prompt, you wait, something arrives. Sometimes it is brilliant. Sometimes it is wrong in an interesting way. Sometimes it is wrong in a boring way. The loop is short: five minutes, sixty minutes, the cache windows you are unconsciously syncing to and the reward is highly variable. That is the textbook recipe for what they call "compulsive engagement". It is not an accident that you keep going past the point where you should stop. The loop is doing exactly what loops like this do. > > The pull you feel toward "just one more iteration" is not evidence that the work needs more time. It is evidence that the schedule of small wins has trained your brain to expect another one. Video games and slot machines feel productive in exactly the same way. The difference is that with AI, sometimes you really did just ship something which makes the pattern harder to see and harder to break. > > Call a savepoint. Have the agent capture the current state. All your open threads in a structured note and then close the conversation. Walk away. When you come back, fresh model and fresh you, you load the savepoint and continue. You can call for a savepoint at any time. You need one and so does your coding assistant. The work will be there when you get back. *Originally posted on [LinkedIn](https://www.linkedin.com/posts/jesserobbins_hey-feeling-depleted-after-working-with-activity-7458153361582751744-9PbM) on 2026-05-07.* ### Musical AI Raises $4.5M: Jesse Robbins Joins Board - Los Angeles Times · Article · 2026-02-04 - Original: https://www.latimes.com/b2b/entertainment/story/2026-02-04/musical-ai-raises-4-5m-funding-heavybit - Canonical: https://jesserobbins.com/mentions/musical-ai-raises-4-5m-jesse-robbins-board-la-times/ Los Angeles Times follow-on coverage of Musical AI's $4.5M round led by Heavybit, with me joining the board. The piece frames Musical AI as rights-aware infrastructure for generative music. ### What I look for when I invest - Funded Podcast (airCFO) · Podcast · 2026-01-28 - Original: https://www.aircfo.com/resources/funded-why-market-size-trumps-everything-in-vc-deals-jesse-robbins-heavybit - Canonical: https://jesserobbins.com/mentions/what-vcs-look-for-ai-investment-founder-tips-market-size/ On airCFO's Funded podcast I walked through how I evaluate companies at Heavybit, why market size is the first filter I apply, and how I read co-founder dynamics before I read the pitch. Alex Wittenberg ran a long conversation with me on airCFO's Funded podcast covering my personal and current firm investment criteria, how to get on a VC radar, and the founder mistakes I see killing pitches before they start. ### Leverage and Focus The biggest advantage early-stage founders have is speed. Anyone on the team can make a decision and ship it the same day, without a process or an approval chain to clear. A good plan executed immediately beats a perfect strategy debated for months. The founders who spend that advantage on real customer work are the ones who build real businesses. The ones who spend it on optics like big-name partnerships do not. "We really, really, really like to see traction: early, repeatable, measurable, personal, where you can say, this is why this customer is using it." ### The Big Lie About Investment Stages There is a big lie in venture: the published stage gates that say you can raise a Series A at $500K ARR. Those numbers are the floor for getting a coffee meeting. The bar for a term sheet sits much higher, and founders who plan against the low end set themselves up for heartbreak. Pre-seed investing means testing a concept with referenceable design partners at zero revenue. Seed means proving there is a repeatable business, typically $500K to $1.5M ARR. AI is compressing how fast founders reach those numbers, but the bar for a term sheet has not moved with it. ### Market Size Is the First Filter Of the four buckets we evaluate (team, market, product, traction) market is the one I will not bend on. A great team with real traction in a small market is still a pass for us. The cleanest example I have is LaunchDarkly. Every VC dismissed feature flagging as too small a market, and Edith Harbaugh reframed it from a developer tool to a business configuration platform. Same product, bigger market, and the round happened. "We can love you. We can believe you are building the best product imaginable. We can see early traction that's super compelling. And if it's not a big enough market, it's not a venture scale business." ### Co-Founder Dynamics I pay close attention to how co-founders treat each other in unguarded moments. The Gottman frame of "turning towards" responses is the one I use: even in conflict, are the founders still in each other's corner? The Sanity founding team (Magnus, Simon, and the rest of the co-founders) has worked together for over a decade and visibly likes working together. That is the version that holds up when things get hard. The inverse is devastating: "When founder relationships fall apart, it is the most value destroying, destructive thing that can happen to you professionally." ### Founder Grit The trait I care about above everything else is personal grit. The ability to grind in the absence of positive feedback, on purpose, because the work matters to you. I trained as a firefighter, and when someone calls 911 the team shows up and solves the problem no matter how big it gets. That is the same job description for a founder CEO. It does not require a Type A personality. Mitch Hill, the CEO I hired at Chef, was a quiet introvert who had previously scaled a company from zero to a billion in revenue, and his resolve was extraordinary. ### How to Get on VCs Radar The playbook: attend startup and VC events, engage with the community, get a warm introduction, or fill out the application form. What does not work: cold LinkedIn requests, AI-generated SDR emails, and launching into an elevator pitch the moment you meet a partner. "All I need to believe at first is that you're the kind of person I want to talk to for an hour or more." ### The Investment Process My personal standard is that every founder should leave thinking it was one of the best VC meetings they have ever had, regardless of outcome. They should actually benefit from the meeting. The pet peeves that kill momentum for me: view-only pitch decks (send downloadable artifacts), confusing SAFEs with term sheets, and trying to manufacture urgency with fabricated interest. ### This is an Adventure! A personal mantra "a journey with certain adversity, uncertain outcome, and good companionship." Transcript: https://jesserobbins.com/mentions/what-vcs-look-for-ai-investment-founder-tips-market-size.md ### Musical AI Bags $4.5M to Scale AI Attribution Tech: Jesse Robbins Joins Board - Music Business Worldwide · Article · 2026-01-13 - Original: https://www.musicbusinessworldwide.com/musical-ai-bags-4-5m-in-funding-round-to-scale-ai-attribution-tech/ - Canonical: https://jesserobbins.com/mentions/musical-ai-funding-round-mbw-jesse-robbins/ > "Musical AI's attribution technology is essential infrastructure that will enable and accelerate every media-focused AI product." > — Jesse Robbins Music Business Worldwide broke Musical AI's $4.5M round, led by Heavybit, with me joining the board. The piece covers what Musical AI is building and why attribution matters for generative music. ### Investing in Vibrant Labs: AI Agent Simulation Infrastructure - Heavybit · Article · 2025-12-03 - Original: https://www.heavybit.com/press/heavybit-welcomes-new-member-vibrant-labs - Canonical: https://jesserobbins.com/mentions/investing-in-vibrant-labs-ai-agent-simulation/ > "Vibrant Labs opens a new frontier in AI infrastructure: production-grade, RL-ready simulation and verifier-driven evaluation built for long-horizon agents." > — Jesse Robbins Annoucing investment in Vibrant Labs, which builds production-grade simulation and verifier-driven evaluation for long-horizon AI agents. I wrote the Heavybit announcement when Vibrant Labs joined the portfolio. Shahul ES and Jithin James, the team behind [Ragas.io](https://www.ragas.io), are building production-grade simulation environments and verifier-driven evaluation for long-horizon AI agents. The operational gap is specific. Agents take on multi-step planning and execution, and developers need safe, repeatable environments to measure and improve agent behavior before deployment. Vibrant Labs ships three pieces of infrastructure: - RL-ready simulation worlds for training agents against realistic scenarios. - Verifier-driven evaluation that returns explanatory feedback alongside pass-or-fail signal. - Training pipelines that turn evaluation data back into supervision signal. > "Vibrant Labs opens a new frontier in AI infrastructure: production-grade, RL-ready simulation and verifier-driven evaluation built for long-horizon agents." ### What Investors Look For in AI Startups: Builders with Taste - Shift Conference · Video · 2025-06-12 - Original: https://www.youtube.com/watch?v=bIw2pvA3Z5k - Canonical: https://jesserobbins.com/mentions/what-investors-look-for-ai-startups-shift/ At Shift Conference Miami 2025 I walked through what makes a developer-tools startup investable, why AI is still in the toil-automation phase, and where things are headed. At Shift Conference Miami 2025, Marco interviewed me on stage in a rapid-fire format covering what makes a developer-tools startup investable, where AI disruption is actually happening in developer workflows, and why the arc of software engineering keeps bending toward natural language. ### Test and Taste My personal investment filter is visceral: does the product change how you think about the problem from the very first time you see it? Two portfolio examples from my time at Heavybit: Tailscale, after building VPNs throughout my career, the first time I used it "made me mad. It was so easy." And Continue, which made me switch from vi to VS Code, "which is sort of a crazy shift for an old school systems engineer." The broader criteria I use: builders with great taste, desperate to solve a problem they need to fix in the world, who can get early developer communities excited. Open-source software with a sophisticated business model. Tools that are hard to build without specialized knowledge. ### Still Automating Toil We are at the very beginning of this. Code generation and basic test automation are low-hanging fruit: "the work that none of us ever really want to do in the first place. That's what we're automating right now. We're automating toil." The real opportunity lies ahead: observability, CI/CD workflows, UI/UX generation on the fly, customization, personalization. Every category of developer tools is getting changed, but the surface covered so far is small. ### The Next Abstraction Layer I think about AI code generation as the latest step in computing's longest arc: from hand-wired circuits to assembly language to increasingly natural programming languages. Each layer of abstraction moved humans closer to expressing software concepts in natural language, with 10x–100x productivity gains each time. "We've arrived at the point where we're able to express software concepts and have another layer of abstraction that generates code, just like every cycle prior to that point." The result is that more people can build software, which creates more complexity to manage, and more developer tools to invest in. "Every time we add a few million more people to building things, well, they get more complicated. There's more of it to manage. And so there's more things for me to build and invest in." ### Agents Are Just Another Developer The reframe I keep coming back to: "Agents are just another type of developer." They need the same things human developers and previous "non-human users" like Google's crawler needed: good documentation, clearly defined APIs, and affordances that make interfaces easy to program against. There is a direct line from accessibility (screen readers) to SEO (Google's crawlers) to agents: each is a non-human user that benefits from the same infrastructure investments. The agentic layer builds on prior work and drives personalization, dynamic UI, and adaptive experiences. "Does it mean that any of the work we were doing before goes away? No. Hopefully it just causes people to do more." Transcript: https://jesserobbins.com/mentions/what-investors-look-for-ai-startups-shift.md ### The Future of Dev Tools Is Autonomous: Engineers Will Become Fleet Generals - Shift Magazine · Article · 2025-05-22 - Original: https://shiftmag.dev/deveeloper-tools-ai-software-engineering-5299/ - Canonical: https://jesserobbins.com/mentions/future-devtools-autonomous-fleet-generals-shiftmag/ Shift Magazine surveys autonomous AI agents in developer workflows and quotes me from the Shift Miami panel on designing software for agents as much as humans. Marin Pavelić at Shift Magazine surveys how developer experience is evolving from desktop-era IDEs to AI-collaborative environments where engineers increasingly manage fleets of autonomous agents. The piece walks through the category (Cursor valued at $9 billion, Windsurf acquired by OpenAI for $3 billion) and asks what happens when AI operates autonomously inside development workflows. ### Designing for agents Marin pulled three lines from the Shift Miami panel "Investing in Dev Tools in the Age of AI." The first: > "If you're building software now, you're not just designing for humans. You're designing for agents, too." This reframes developer experience as a dual-audience problem. The tools, APIs, and documentation we build now serve both human developers and the AI agents working alongside them. It echoes the longer argument from my [Shift Conference interview](/mentions/what-investors-look-for-ai-startups-shift/) that agents are "just another type of developer" who need the same affordances as human users. ### Open source reignited my joy The second line is personal: > "Because of this open-source ecosystem, I started writing code again. It felt joyful." The tools finally feel right. I came back to hands-on coding because Continue and the AI-native open-source community made it worth doing. ### Collaboration over fighting The third line is the one I keep returning to: > "Experiencing joy in collaborating with tools instead of fighting them may be the most important change." The article places this alongside the broader frame I have been making everywhere: delegation as the new automation, engineers as fleet generals orchestrating AI agents, and open-source tools like Continue providing the transparency and control that make the collaboration possible. ### Experimentation, Causal Inference, and AI: Sean Taylor of OpenAI and VC Jesse Robbins at Data Council 2025 - Heavybit · Article · 2025-04-01 - Original: https://www.heavybit.com/library/article/data-council-2025-the-data-science-and-algorithms-track-with-sean-taylor-and-jesse-robbins - Canonical: https://jesserobbins.com/mentions/data-council-2025-sean-taylor-openai-jesse-robbins-investor/ > "AI provides an opportunity to radically improve how we do things." > — Jesse Robbins I interviewed Sean Taylor of OpenAI ahead of Data Council 2025 on experimentation, causal inference, and why AI generates more questions that need empirical answers. I sat down with Sean Taylor of OpenAI at Heavybit to preview the Data Science and Algorithms track for Data Council 2025. Our conversation focused on experimentation, causal inference, and the practical frameworks data teams use to drive product and business decisions. We covered the featured speakers: Hadley Wickham on generative AI in data science workflows, Timothy Chan of Statsig on experimentation at scale, Joe Powers of Intuit on Bayesian A/B testing, and Bryan Bischof of Hex on ML engineering. I framed the conference as a rare opportunity for isolated data science practitioners to validate their work by "peeking over each other's shoulders." ### Next in Tech Ep. 197: Data Pipelines for AI - S&P Global Market Intelligence · Podcast · 2024-12-10 - Original: https://www.spglobal.com/market-intelligence/en/news-insights/podcasts/next-in-tech-ep-197-data-pipelines-for-ai - Canonical: https://jesserobbins.com/mentions/data-pipelines-for-ai-next-in-tech-podcast/ > "Data pipelines are having a DevOps moment, starting with a cultural and technical shift toward continuous integration and delivery." > — Jesse Robbins On S&P Global's Next in Tech, Eric Hanselman and Jesse Robbins discuss enterprise AI data pipelines and why infrastructure quality matters more than model scale. Eric Hanselman hosted me on S&P Global's Next in Tech to argue that enterprise AI success depends on data pipeline quality more than model scale, and that the infrastructure patterns for getting it right already exist in the DevOps playbook. ### The DevOps moment for data Enterprise AI data infrastructure is in the same place software delivery was a decade ago. DevOps introduced continuous integration and delivery to replace brittle manual deployment, and data teams are now building pipeline automation for model training, evaluation, and governance. The cultural shift matters as much as the technical one. Organizations have to treat data delivery with the same rigor they eventually brought to code delivery. ### Start small, iterate locally I advocate starting with smaller models and localized datasets. The approach is the same one I have run since Chef and Amazon: prove the pattern works at small scale, measure what matters, then expand. Enterprises that try to solve the data pipeline problem at full scale first build fragile architectures and burn through budgets before they learn what actually works. ### Pipeline quality as competitive advantage The companies that win at enterprise AI are the ones that build the best data pipelines. Clean data in. Reliable inference out. Governance and cost controls at every stage. The data delivery layer is where the next generation of developer tools and infrastructure companies will be built. ### The Data Pipeline is the New Secret Sauce - Heavybit · Article · 2024-09-16 - Original: https://www.heavybit.com/library/article/ai-infrastructure-top-challenges-data-inference - Canonical: https://jesserobbins.com/mentions/the-data-pipeline-is-the-new-secret-sauce/ > "The biggest challenge emerging is building and operating the infrastructure both for creating and running the data pipelines to build, manage, and maintain a robust, secure body of proprietary data." > — Jesse Robbins I wrote this at Heavybit in September 2024. The argument: the data pipeline is the differentiating asset in enterprise AI. Includes four inference hosting models and four enterprise maturity phases. The data pipeline is the bottleneck in enterprise AI. I wrote that at Heavybit in September 2024. The argument: enterprises that will get real work out of AI are the ones whose pipelines produce a secure first-party dataset and stay correct as the underlying systems change. Buying access to a model is one part of shipping AI into production. The harder work is operational. Heavybit Library article, September 2024, arguing the data pipeline is the differentiating asset in enterprise AI and mapping four inference hosting models against four maturity phases. In September 2024, roughly 40% of enterprises surveyed said they had deployed an AI program or were actively exploring one. Microsoft was reporting 53,000 organizations using its AI offerings via Azure. Gartner had 87% of "mature organizations" carrying dedicated AI teams. These programs are not bought off the shelf. They produce an artifact: the internal dataset. It is the end result of a complicated toolchain run by a team that does this work full time. The data pipeline is that artifact's production line. It is the thing that does or does not compound. The piece maps four inference hosting models and four enterprise maturity phases. Most of the value is in the phase model. Phase 1 looks like real experimentation against a hosted API. Phase 2 is the moment teams realize they have to stand up an internal pipeline to extract durable value from the use cases they have proven. Phase 3 is cost shock. The bill from the API provider becomes the line item that gets the AI program in front of the CFO. Phase 4 is the only one that requires judgment, because the right answer for one workload is rarely the right answer for another. Mature enterprises end up running mixed inference configurations and treating optionality as the durable asset, not any single hosting choice. The operational point underneath the taxonomies is the one I care about. A data pipeline is a continuous process that begins when the model ships to production. It needs the same monitoring, validation, and team discipline as any other production software, plus the security and privacy work that stops personally identifiable information from leaking on the way through. Without those practices, enterprises either fail to build their internal dataset at all or ship real business risk: privacy leaks, model drift, the cost of re-training on bad data. Cost-control work belongs in this phase too, including model merging and mixture-of-experts as alternatives to retraining on the entire dataset. In 2025, Heavybit backed Recce. I joined the board. CL Kao and his team are building data validation for the moment a pipeline change ships to production, which is exactly the discipline this article said enterprises would have to develop. Recce is the practical answer to a question this piece could only frame: how do you know your pipeline is still correct after the change? > The biggest challenge emerging is building and operating the infrastructure both for creating and running the data pipelines to build, manage, and maintain a robust, secure body of proprietary data to train, fine-tune, and orchestrate LLM operations, and for running inference, the actual process of models running calculations on inputted data. The piece runs through two frameworks. **Four inference hosting models.** - **Hosted API.** Calling [OpenAI](https://openai.com/), [Anthropic](https://www.anthropic.com/), and the rest. The provider absorbs the cost and operational burden of running large models. Enterprises pay in tokens. - **On-device edge.** Smaller models running locally, often on high-end laptops, sometimes paired with three-billion-parameter open-weight checkpoints. Lower latency, better data locality, an unclear scaling story for larger teams. - **On-premise data center.** Everything behind the firewall. Most enterprise IT workloads moved off-prem years ago for total-cost reasons. AI inference is following the same path outside heavily regulated workloads. - **Off-premise cloud via third-party data center.** The managed-inference layer. Resembles traditional cloud computing more every quarter. Introduces network latency and dependency on the provider's reliability posture. **Four enterprise maturity phases.** - **Phase 1, off-the-shelf cloud start.** Most enterprises begin here, against a hosted API. Data science and operations teams focus on identifying valuable use cases. The hosting question is abstracted by the provider contract. - **Phase 2, scaling what works.** Teams have a pipeline that is good enough to deliver value on specific workloads. They harden privacy posture for the data types and jobs that matter. The bill starts to register. - **Phase 3, cost shock and optimization.** API spend hits a number that gets noticed. Teams reassess: continue paying for hosted inference, or invest in an internal model and the inference configuration to run it. Tooling investments here include pretraining datasets, data filtering, model evaluation. - **Phase 4, specialization.** Mature teams run mixed configurations and stop treating any single hosting choice as the answer. They prize optionality over vendor lock-in. They consider [model merging](https://arxiv.org/abs/2403.13257) and [mixture-of-experts](https://arxiv.org/abs/2305.14705) as alternatives to retraining on the entire dataset every time. The piece names a specific operational point. The pipeline itself is the artifact, and building one demands operational effectiveness plus the security and privacy work that stops PII from leaking. Without that, enterprises fail to build their internal dataset at best, or ship real business risk from privacy leaks, poor model performance, and the cost of retraining on bad data. The article also references [Chaoyu Yang](https://www.linkedin.com/in/chaoyuyang/) of [BentoML](https://www.bentoml.com/) on specialized AI systems built for specific use cases as a likely source of durable advantage. Specialized infrastructure is where serious teams have ended up in 2025 and 2026. *Read the full article at [Heavybit](https://www.heavybit.com/library/article/ai-infrastructure-top-challenges-data-inference).* ### AI Investor Jesse Robbins on NYSE Floor Talk - NYSE · Video · 2024-08-05 - Original: https://www.youtube.com/watch?v=nK6pEkjv7tU - Canonical: https://jesserobbins.com/mentions/jesse-robbins-nyse-floor-talk/ > "I am focused on investing in pre-seed and seed companies using AI to enable new ways of writing software, of managing and deploying the software and infrastructure that powers everything." > — Jesse Robbins On NYSE Floor Talk, August 2024: a short statement of what I invest in now, AI-powered developer tools and infrastructure at the pre-seed and seed stage. ## Full Interview Transcript I invest in early stage AI companies focused on developer tools, productivity, and infrastructure. I like to think I invest in companies that power the companies that you hear of every day. **Q: Jesse, as you look to the next 12 months what are your priorities right now?** I'm focused on investing in pre-seed and seed companies, often where there is a strong AI component to enabling new ways of writing software, of managing and deploying software, of managing infrastructure that kind of powers everything. The last 12 months have shown us there's this incredible shift occurring, and so the next 12 months is going to be about helping our companies scale and grow and hire and raise more Capital and to reach a lot more Enterprise companies that are our amazing customers. **Q: Jesse tell me who are some of your companies?** Some of our larger late stage companies or companies like PagerDuty that was listed on NYSE. Other ones like LaunchDarkly and Snyk and Sanity, as well as smaller companies that are sort of at the pre-seed and Seed stage including Shipyard and Radar (a New York company), Mobot (another New York company), and many others. ### Jesse Robbins Named One of the 30 Most Successful Early-Stage Startup Investors - Business Insider · Article · 2024-01-17 - Original: https://www.businessinsider.com/30-of-the-most-successful-early-stage-investors#jesse-robbins-heavybit-16 - Canonical: https://jesserobbins.com/mentions/jesse-robbins-named-top-30-early-seed-stage-vc-investor-business-insider/ Business Insider named Jesse Robbins one of the 30 most successful early-stage startup investors of 2024. An investor in developer tools and infrastructure, his entry cited investments in Fastly, PagerDuty, LaunchDarkly, and CircleCI. Robbins cofounded Chef and the DevOps movement. Business Insider named me to its 2024 list of 30 most successful early-stage startup investors. Grateful to the founders, the practitioner community Velocity and Chef were built with, and my partners at Heavybit. ### Heavybit Welcomes New Member: Continue - Heavybit · Article · 2023-11-14 - Original: https://www.heavybit.com/press/heavybit-welcomes-new-member-continue - Canonical: https://jesserobbins.com/mentions/heavybit-welcomes-new-member-continue/ > "I'm excited to welcome our newest portfolio company, Continue, which gives software engineers the power to streamline their development process using large language models (LLMs) and hit flow state faster and longer." > — Jesse Robbins Heavybit's announcement when Continue joined the portfolio, an open-source tool that brings LLM assistance directly into the IDE. I wrote this announcement when Continue joined the Heavybit portfolio. Continue integrates large language models directly into VS Code and JetBrains, so engineers do not have to copy code between applications, make edits, and re-prompt separate AI tools. The reason we backed the team is the IDE-native, open-source design. Developer tools succeed when they meet engineers where they already work. Continue does that. ### Cloud Native StartupFest 2023 - CNCF / KubeCon · Panel · 2023-11-06 - Original: https://www.youtube.com/playlist?list=PLj6h78yzYM2NDTB17J_VLemCUIJlfbHsM - Canonical: https://jesserobbins.com/mentions/cloud-native-startupfest-2023-kubecon-jesse-robbins/ > "Open source is not a business model. Open source is a movement. We're still figuring out the business models." > — Jesse Robbins I co-hosted Cloud Native StartupFest at KubeCon NA 2023 with Erica Brescia and Dave Zilberman: fundraising in the post-2022 capital environment, open source business models, and what investors actually look for. The full transcript of my opening remarks is below. Erica's data slides and the founder panels are on the [CNCF YouTube playlist](https://www.youtube.com/playlist?list=PLj6h78yzYM2NDTB17J_VLemCUIJlfbHsM). Transcript: https://jesserobbins.com/mentions/cloud-native-startupfest-2023-kubecon-jesse-robbins.md ### Generative AI in DevOps and Incident Response: What the Experts Actually Think - Heavybit · Article · 2023-10-12 - Original: https://www.heavybit.com/library/article/generative-ai-incident-response-devops - Canonical: https://jesserobbins.com/mentions/2023-10-12-generative-ai-devops-incident-response-heavybit/ > "GenAI is good at confidently delivering text that is pleasant to read, but not always complete, or correct." I interviewed Nora Jones, Jeremy Edberg, Mandi Walls, and Brent Chapman on what generative AI actually does in incident response, and where humans have to stay in the loop. By late 2023, every vendor had an AI pitch for incident response. Autonomous remediation. AI-powered root cause analysis. MTTR cut in half, automatically. I went to the people who actually run systems at scale to find out what was real. I interviewed [Nora Jones](https://www.linkedin.com/in/nora-jones-0b3b7b1a/) (founder of [Jeli](https://www.jeli.io), formerly Slack and Netflix), [Jeremy Edberg](https://www.linkedin.com/in/jedberg/) (Amazon Alexa, formerly Netflix and Reddit), [Mandi Walls](https://www.linkedin.com/in/mandiwalls/) ([PagerDuty](https://www.pagerduty.com), formerly Chef), and [Brent Chapman](https://www.linkedin.com/in/brentchapman/) (Great Circle Associates, formerly Google and Slack). What came back was more specific, and more cautionary, than the vendor hype. ## Key Themes ### Where AI actually helps: summarization The consensus is narrow but real. AI is useful for incident summarization: drafting post-incident reports, generating status updates, and catching latecomers up to speed during an active incident. These are tasks where a plausible first draft is valuable, and where a human will verify before it matters. Nora Jones frames it precisely: "I think what we really want to do is use AI to get people more curious about what's happening in their incidents." The value is in the question it opens, not the answer it provides. ### Why Hallucination Disqualifies Autonomous Remediation The harder truth is that AI hallucinates, and in incident response, hallucination is dangerous. Brent Chapman's observation cuts to the bone: "LLMs are sometimes wrong, but never uncertain." A system that is confident and occasionally wrong is the worst possible profile for taking autonomous action in high-stakes, time-compressed situations. Jeremy Edberg draws the line explicitly: "Right now, GenAI is something of an advisory tool. We're not to the point where we trust it enough to take actions." The human-in-the-loop is not a temporary limitation waiting to be engineered away. It is the right architecture given current reliability. Every AI-powered RCA and autonomous remediation pitch in 2023 ran into this constraint. The serious practitioners all landed in the same place. ### The Junior Developer Analogy Multiple practitioners reach for the same frame independently: working with AI tools is like working with a very junior programmer. You still have to check everything. You still have to ask whether the output is right, useful, and complete. Alert correlation, automated runbooks, AI-assisted postmortem generation all require the same oversight. Edberg uses this frame to address career anxiety directly: "If you are good at logic and want to learn how to reason about computer systems, software engineering will still be a great place to be." AI will accelerate development, particularly for engineers who already know how to think about systems. It will not replace that judgment. ### The AI SRE Career Question Mandi Walls draws a useful distinction for roles: SREs, who are deeply integrated with engineering practice, will see more direct benefits from AI coding and observability tools than operators in more traditional roles. The productivity gains land where the work is closest to the code. MTTR reduction through AI-assisted triage and automated runbooks accrues to engineers who understand what the AI is actually doing. The overall career picture is a shift toward strategic work and away from the mechanical. Practitioners who understand systems reasoning will find AI accelerates their output. Those optimizing for task execution without deeper systems understanding face more pressure. ### The Human Judgment Floor What runs through every perspective is a shared conviction: AI does not yet have contextual or collaborative judgment. It cannot read the room during an incident, weigh the organizational history that shapes which escalation matters, or decide what to prioritize when the situation is genuinely novel. Autonomous incident investigation fails at exactly the moments when the incident is most unusual, which are precisely the moments when it would matter most. Nora Jones is direct: "I don't think generative AI is going to fix the incidents for you." Someone still has to verify, lead, and decide. The tools make some of that work faster while keeping the responsibility in human hands. ## Further Reading - [What to Know About the Modern Incident Response Lifecycle](/mentions/incident-response-best-practices-heavybit/) — Earlier Heavybit piece on incident response fundamentals, also featuring leading practitioners - [Resilience Engineering: Learning to Embrace Failure](/mentions/resilience-engineering-learning-embrace-failure-acm-queue/) — Jesse's ACM Queue contribution on resilience engineering principles that underpin modern incident response - [GameDay: Creating Resiliency Through Destruction](/mentions/gameday-creating-resiliency-through-destruction-usenix/) — Jesse's USENIX LISA talk on the GameDay exercises he started at Amazon, a precursor to chaos engineering - [Fireside Chat with Jesse Robbins and Kolton Andrus — Failover Conf 2021](/mentions/fireside-chat-jesse-robbins-kolton-andrus-failover-conf/) — Jesse and [Kolton Andrus](https://www.linkedin.com/in/koltonandrus/) on chaos engineering and resilience a decade after GameDay ### DevOps is dead? Nope, it is maturing ft. Jesse Robbins - The Confident Commit · Video · 2023-04-07 - Original: https://podcasts.apple.com/us/podcast/devops-is-dead-nope-it-is-maturing-ft-jesse-robbins/id1565433605?i=1000607861168&uo=4 - Canonical: https://jesserobbins.com/mentions/devops-is-dead-nope-it-is-maturing-confident-commit-podcast/ > "Organizations evolve like cities. You start with a few shacks in the woods. Eventually you have enough at stake that you need building codes, fire codes, a fire department, and someone who actually tests the sprinklers." > — Jesse Robbins DevOps is not dead. It's maturing. Platform engineering is the next layer of the same idea, not a replacement for it. My conversation with Rob Zuber on what's actually changing and what isn't. Transcript: https://jesserobbins.com/mentions/devops-is-dead-nope-it-is-maturing-confident-commit-podcast.md ### What to Know About the Modern Incident Response Lifecycle - Heavybit · Article · 2022-11-11 - Original: https://www.heavybit.com/library/article/incident-response-best-practices - Canonical: https://jesserobbins.com/mentions/incident-response-best-practices-heavybit/ > "Teams only get good at this when they embrace the whole process and each of its steps." > — Jesse Robbins Heavybit's incident management guide quotes me on why teams only get good at incident response when they treat the whole lifecycle as one discipline. Andrew Park's guide for Heavybit on modern incident management quotes me on why teams only get good at this when they treat the full lifecycle as a single discipline. The piece walks readers through the practical implications: normalize incidents by talking about them often, be honest about the state of the infrastructure, and treat the practice itself as the source of mastery. The line he pulled from our conversation is the one I have been saying since Amazon: skip any step in the cycle and you never fully develop the muscle for any of them. ### Fireside Chat with Jesse Robbins and Kolton Andrus • Failover Conf 2021 - Gremlin · Video · 2021-04-29 - Original: https://www.youtube.com/watch?v=6E_caMdCDIY - Canonical: https://jesserobbins.com/mentions/fireside-chat-jesse-robbins-kolton-andrus-failover-conf/ At Gremlin's Failover Conf 2021, Kolton Andrus and I covered GameDay origins at Amazon, the evolution of chaos engineering, and where reliability practices were headed. At Gremlin's Failover Conf 2021, Kolton Andrus and I sat down for a fireside chat on GameDay's origins at Amazon, how deliberate failure injection evolved into the discipline the industry now calls chaos engineering, and where reliability practices needed to go next. We covered the early days of breaking production systems on purpose, the cultural resistance you hit when you first propose simulating catastrophic failures, and how those exercises changed the way Amazon thought about availability. We traced the shift from ad-hoc failure testing to systematic chaos engineering platforms, and dug into what separates teams that recover well from incidents from teams that struggle. The session closes on the future of SRE: what engineering leaders should prioritize and how the chaos engineering community can continue raising the bar on production resilience. ### The Seed 100: The Best Early-Stage Investors of 2021 - Business Insider · Article · 2021-04-02 - Original: https://www.businessinsider.com/seed-100-top-early-stage-vc-investors-2021-4 - Canonical: https://jesserobbins.com/mentions/seed-100-businessinsider-2021/ > "Robbins is the right investor to call in an emergency." > — Business Insider Business Insider named Jesse Robbins to its 2021 Seed 100, its ranking of the best early-stage venture investors, built with Tribe Capital. An investor in developer tools and infrastructure, his entry cited seed investments in Conjur, LaunchDarkly, and Zymergen. Robbins cofounded Chef and the DevOps movement. Business Insider named me to the 2021 Seed 100, a ranking of early-stage investors built with Tribe Capital from data on about a thousand seed investors. ### An oral history of #hugops: How tech's first responders built a culture of empathy - Protocol · Article · 2021-02-25 - Original: https://www.protocol.com/enterprise/oral-history-hugops - Canonical: https://jesserobbins.com/mentions/oral-history-hugops-protocol/ > "I've got to change the way that I approach this entirely and make it safe to experiment." > — Jesse Robbins Protocol's oral history of

A note from Jesse

Tom Krazit wrote this oral history of #hugops for Protocol in February 2021. Protocol shut down the following year, so I am preserving it here.
*When something breaks on the internet, the people who know how to fix it just want to give their colleagues a hug, even if they're a rival.* ![The #hugops community in its happy place: the Velocity conference in San Jose, 2010](/images/mentions/hugops/velocity-2010-crowd.jpg) *The #hugops community in its happy place: the Velocity conference. Photo: James Duncan Davidson/O'Reilly Conferences* --- In almost every profession, it seems like there are two types of workers: the ones who get the glory, and the ones who do the essential work no one ever sees, unless something goes wrong. In enterprise computing, those overlooked people are known as operations engineers. They're the ones who keep the rickety Rube Goldberg machine that is the modern internet from falling to pieces every day, while their glamorous counterparts, software developers, get to bask in the recognition that comes with shipping a new feature or creating a new service. A little over 10 years ago, a group of operations-oriented engineers decided they were fed up with software developers who didn't care if their code actually worked, so long as it shipped. They were tired of abuse at the hands of management who forced their teams to be on call 24/7 with little to no internal support, let alone recognition. Those engineers created the Velocity Conference in order to band together: to share their lived experiences including the intense pressure to keep Fortune 500 companies up and running, to discuss tips and tricks for navigating tricky problems and to come together as a community of people who know what it's like to be at the bottom of the food chain when everything has gone to hell. That community sparked a revolution known as DevOps, the idea that software developers and operations professionals needed to work together more closely to support the ever-more complex task of running sophisticated software over the internet. Big companies such as Amazon and Google started to develop the operations career path with incentives and rewards parallel to those on the development side, while acknowledging that these people needed support from the highest levels of the company to do their very difficult jobs. And out of this community came a Twitter hashtag, an in-group signal to their peers during the most stressful moments of their careers that a team had their back. When a major cloud service goes down, such as during Slack's early January outage, most people on Twitter see an opportunity to vent their frustration and score points at the affected company's expense. At those moments, the people who know what it takes to keep these services afloat spread a hashtag: #hugops. This is the story of the engineers who keep the cloud running, and how they created their own culture of empathy when nobody else cared. ## Life of a sysadmin **Adam Jacob**, CEO of The System Initiative, co-founder and former CTO of Chef: Systems administrators, a now almost basically nonexistent job title, were not the most beloved humans in the technical world. We didn't get a lot of respect. We were sort of in the same bucket like secretaries; we had a System Administrator Appreciation Day. The people who do the stuff you don't see get appreciation days because, by definition, it means I'm not being appreciated every other day. [One team leader] took us and my whole team, there were like 20 systems administrators, and he took us all out for beer on System Administrator Appreciation Day. And he sat down with the pitchers of beer and the first thing he said was, "Here's your guys' beer. Too bad none of you are smart enough to be engineers. Cheers." My response to that was to just be mean to him. **Jennifer Davis**, developer relations manager, Google: I don't know if you ever heard of the BOFH sysadmin? There was this mentality of like, how cruel and evil can we be to our users. **Werner Vogels**, CTO, Amazon: I think sysadmins mostly came out at a time when most companies were buying software. Traditionally at those operations, [software] development is on one side. Then there's this wall, and you throw software over the wall; and you don't care anymore. **Tim O'Reilly**, founder, O'Reilly Media: There were all the, effectively, software janitors who were cleaning up after them. And the software janitors were kind of going: That doesn't really work. ![Jesse Robbins on stage at the Velocity conference in 2010](/images/mentions/hugops/jesse-robbins-velocity-2010.jpg) *Jesse Robbins, a former firefighter and present-day hugger, at the Velocity conference in 2010. Photo: James Duncan Davidson/O'Reilly Conferences* **Kolton Andrus**, co-founder and CEO, Gremlin: At Amazon, I was one of 10 people that was paged when the website went down. And I took and managed the resolution of those calls from the side of the freeway next to my motorcycle because I had to pull over, call in and handle it immediately; it couldn't wait 10 minutes until I got home. There was an Amazon Christmas party that I was at where I got a page, I had to run out to my car, get my backpack, come into a war room, sit down and resolve an incident before going back to the party. There's a lot of work that the engineers and the ops folks do behind the scenes, a lot of thankless work to help make sure things go well and get fixed. **Nathen Harvey**, developer advocate, Google: What do we celebrate in technology? We celebrate new; new features, shipping new capabilities that we're delivering to customers. And we get angry when systems fail. Basically what you're saying is: We celebrate the developers, and we recognize the operators when everything goes to shit. That's not great. **Jacob**: I sat in a room early on at Chef with a bunch of video game developers that were running the U.S. operations for one of the biggest video games of all time. And their boss sat across the table from them, and to my face, in front of them, said, "My guys aren't smart enough to learn Ruby." If you just interviewed system administrators from that era, 100% of them have that story. **Jesse Robbins**, founder and executive chairman of Orion Labs, former co-founder and CEO of Chef: In operations, we always missed the launch party, because we were too busy in the data center or locked in an office looking at green screens trying to support a launch. We were never there for the fun part. We were always the ones that were giving up our nights and our weekends, and we're powerless to actually improve things. ## When emergencies are a day job **Andrus**: The on-call training I received at every company amounted to: "Here's your pager, good luck. You're smart, you'll figure it out." **Harvey**: I remember a conversation I had with Ron Vidal, who is a firefighter in the San Francisco area. And one of the things he said to me was: "A firefighter has never, in their life at work, responded to an emergency. If your house is on fire, that's an emergency for you, but for the firefighters, that's their job." **Robbins**: I'm a firefighter by training, and when I joined Amazon in 2001, "master of disaster" was my title. I realized that the way that we were running operations at Amazon was fundamentally not going to scale and that we needed a process and almost a cultural overhaul. I began turning Amazon into a fire department. I literally took the sort of incident management principles that we used in the fire service and turned that into what we call GameDays and Scale Days, using essentially the incident command system in order to support people through the various ways of thinking when the red light is on. **Davis**: A lot of what operations is like encourages this heroism: You have to do everything to keep it running and just throw yourself into it. It's not sustainable work. It's not great, it's terrible, and you're celebrated when you save the day but the reality is, it's terrible. It harms your relationships, and it harms your health and just frames how you work with other people. **Nora Jones**, founder and CEO, Jeli: We're shifting towards a kind of a time where people see issues and incidents as a symptom rather than a cause of something, and trying to understand the bigger system that is playing out in those organizations. **Robbins**: I owned availability at Amazon, and when I say owned it, I was sort of a tyrant, and ran it very aggressively. There was this big outage that we had [in the early 2000s], and there was a person early in their career who was literally shaking when I walked into the room because they were so afraid of what was going to happen. I realized, "I've got to change the way that I approach this entirely and make it safe to experiment, safe to do these other things, to not have this punitive model and approach." It was seeing that person's face where I'm like, "Oh, I'm not the fire department, I'm like a bad guy. I'm being a villain." **Davis**: If we reduce the heroism, we can reduce burnout. **Robbins**: There is an ethos that came from all of that early work that recognizes how it is important to be kind to each other. And part of what I did early on at Amazon was create a culture of safety. You only get to do really big, great things when you're able to take great risks safely. ## A meeting of like minds **John Allspaw**, founder and principal, Adaptive Capacity Labs: These topics deserved an entire conference. I guess it was less that it deserved an entire conference, but more that a few folks convinced Tim O'Reilly to actually do it. **O'Reilly**: They said, "Look, we need a gathering place for our tribe." We had done that before, for these various open-source communities. A lot of these things are rooted in communities, and so if you can figure out what community you want to bring together, you start by bringing them together. **Allspaw**: What [the Velocity Conference] did was important, because it was a signal that operating software and understanding how things are running and anticipating things that can go wrong could be considered distinct from software development. **Artur Bergman**, co-founder and chief architect, Fastly: What we were doing was just as critical as writing the code. If you can't run the code, it has no value. **Vogels**: The time to develop software is actually quite small [compared] to the time that you have to operate it. So even though you may be building something complex, it may take a year or two years [to build], you may have to operate it for many, many more years to come. **Jacob**: Velocity was like the first time that there was a non-academic place where everybody who is doing that work could get together. And it was like, well-funded and pretty. It wasn't like we were meeting up in the American Legion hall or whatever. It was a fucking conference. **Allspaw**: We were finding this pretty significant common ground. For many, many years, they didn't have a place to put these ideas, or even labels or terms or vocabulary to talk about the dread, or actually sort of outright terror, that can come with, "shit's broken, and we have no idea." So there's this lived experience of, "OK, you're with your colleagues and shit's broken and you don't have 100% clarity, but you've got a couple of good ideas that look sort of fruitful. And okay, so it seems like we should connect this thing to this thing and restart this other thing? We should do it in that order. What do you think about that?" You'd see this in IRC, we didn't have Slack back then. This conference exists because we've got this shared experience with incidents and the general challenge is not just responding to incidents, but trying to work out how to prevent the ones in the future. And it's difficult work. ## Time for a hug **Jacob**: I'm a very huggy person. And so I hugged all of those people [at Velocity], all the time. Because it was happening to this group of people who ... their work environment was not a place where you got a fucking hug. **Davis**: We're building complex systems that include the people. And so how do we handle the unpredictable stress of complex systems? When you think about hugs, hugs are used to reduce pain. They're used to show that you care and they're used to reduce fear. **Jacob**: So Artur Bergman was, is?, a particularly salty dude. He swears as much as I do, maybe more, and he's Swedish, so like when he swears, it's *better*. Artur is not a person who was huggy. Artur would maybe suffer a hug from me, or suffer a hug from John [Allspaw]. At some point, John made a T-shirt that is the earliest I remember of the #hugops-y thing, and on the back of it it basically says, "Hug Artur Bergman." **Bergman**: [During one Velocity] I gave a keynote and then [Adam] gave a keynote where he told people to hug me, and I was not aware that he had said that. During the day around the conference, random people started coming up and hugging me, which was, you know, quite uncomfortable, especially because I had no idea why. And so I ended up hiding for the rest of the day until I finally found out at the end of the day why this was happening. ![Artur Bergman at the Velocity Conference in 2010](/images/mentions/hugops/artur-bergman-velocity-2010.jpg) *Artur Bergman, who is not the naturally huggy type, at the Velocity Conference in 2010. Photo: James Duncan Davidson/O'Reilly Conferences* **Jacob**: It was a very special moment in time where there was this very high degree of camaraderie, there was this really high degree of familiarity. **Allspaw**: Capturing this real dread, these pretty scary, pressure-filled situations, sort of fueled that you're part of this tribe. I don't know who you are, but you're here and you're talking and, so having that common ground is what I think genuinely got people [to be] like, "Can I give you a hug?" **Jacob**: We knew people at all of those [big tech companies], right? And so as everybody starts to know each other, when like, Facebook would have an outage, you'd use the #hugops hashtag and you were like literally talking to your people. **Robbins**: It's not a surprise that what began with a sarcastic joke to troll one of my best friends became an idea that a lot of people have rallied around because it reflects the world that they're building continuously, that they're continuously improving. **Davis**: It's just a message of caring. It's a shorthand to show that I have empathy for where you're at, because I'm going to be there at some point. And I hope you show me that empathy too, but also, you know what? You are not alone. ## The future according to #hugops **Jones**: What we're really seeing right now is a shift in the software industry and us buttoning up and understanding that our software is quite critical. But the pressures that people are under to write this software is a lot. Take Slack. During that outage, they had all just come back, it was the Monday that everyone came back from New Year. I can't imagine being in that office, because you're just getting used to writing code again, you're just getting used to deploying things again, and then all of a sudden, all the world is signing on to Slack at the exact same time. It makes total sense that they had an incident that day. I think part of what we're seeing from the "learning from incidents" community is just a shift in thinking and software to say, "OK, they didn't do something wrong. Something happened that made sense for them to do what they did," and kind of allowing for that conversation to happen. **Robbins**: That shift happened because we made it happen, in part because we simply made it so clear that large businesses, large organizations cannot succeed with this kind of outdated enterprise software legacy mindset. To be always on, to be always available, you're always improving, and that means dealing with failures and enabling rapid change. I think we're in the second chapter now of a movement that has new leaders emerging and evolving. It's not a part of the MBA curriculum yet, but it soon will be. **Andrus**: Inertia within an organization is hard. You can get a team of 10 to pivot quickly. You're a startup, you've got 100 people, you can change your process. You've got 10,000 engineers, it's a lot harder to get everyone to change how they've done things the last decade or two. **Harvey**: The #hugops movement and the ideas behind it really speak about, "How do we build more empathy for the other humans that we interact with every day?" In my mind, it certainly goes beyond technology. As a society, we could take some real lessons from this: How do we just have better empathy for and respect for the work and the way that people show up in the work that they do, and the fact that you know, everyone is out there doing absolutely the best that they can with what they have? I think that's really, really important. **Davis**: Every time I hear "NoOps" or "NoDev," I'm like, "Nooooo...." Because when people are saying that the robots and automation are gonna take over, that doesn't think through all of these complexities that humans are really great at. Yes, reducing the toil is so great. And we can have these conversations about how to balance out what availability is, and how much I'm going to spend on resolving things, and have those kinds of conversations separate from like, "We're gonna just eliminate all the humans because humans make mistakes." Humans make mistakes building the stuff that then we're relying on; we need humans as the safety checks. **Bergman**: If you have a long outage, you need to care about your people and their sleep schedules, and the fact that they have to eat. And by day four or five, if you didn't do that, you're just gonna have a bunch of really tired and grumpy people who are going to make more mistakes. **Andrus**: I did enjoy at Amazon and at Netflix the approach of, "You should know how your software behaves." If you've written software and deployed it and then you're turning a blind eye to it, that's just not good engineering. **Davis**: What is so fascinating is that the next generation isn't putting up with this negative stuff. They're setting the expectations and they're very vocal about what they want their work environments to be like and how they want to work. ![John Allspaw in his #HugOps shirt at the Velocity Conference](/images/mentions/hugops/john-allspaw-velocity.jpg) *John Allspaw values the shared experience of the Velocity Conference. Photo: pinar@pinarozger.com* **Jones**: We need to be asking different questions and we need to give more people seats at the table. I've been at way too many organizations where the incident was just the [site reliability engineers] in the room. It should have had marketing in the room, it should have had PR in the room, it should have had customer service in the room, it should have had leadership in the room. But it's thought of as kind of an SRE issue, like SREs have to prepare for any type of situation that gets thrown their way. I was at one organization a while back where we launched a Super Bowl commercial. And we had some bumps when we launched the commercial, but the SRE team didn't get a ton of notice that the commercial was happening, I think it was either same-day notice or a couple days beforehand, and that was not really mentioned in the post-incident review. **Andrus**: The flip side of #hugops is I do think there is responsibility that should be held to the leadership of those companies. We're empathetic to the engineers that are dealing with the situation they have, but in part that's because leadership isn't prioritizing their actions, or resilience and reliability in the same way that they prioritize some of their product efforts. **Allspaw**: As my colleague Dr. Richard Cook has said, we shouldn't be surprised that these systems go down. We should be more surprised that they stay up as often as they do. **Bergman**: We took a job that was critical to running the world's largest websites and the internet, that was kind of under-appreciated, and turned it into a movement, modernized it with DevOps, and gave those individuals career paths. **Jacob**: Who gets credit when you see a beautiful car? You don't give credit to the mechanics. You're like, "Man, those guys at Porsche really make beautiful cars." You might know, like, one legendary mechanic in the history of great mechanics. But that's why it's so persistent: because the mechanics know the mechanics. ### A Developer's View Into Blockchain Network Architecture - Heavybit · Panel · 2018-08-27 - Original: https://www.heavybit.com/library/article/a-developers-view-into-blockchain-network-architecture - Canonical: https://jesserobbins.com/mentions/a-developers-view-into-blockchain-network-architecture-heavybit/ I joined a Blockdaemon panel at Heavybit with Brian Behlendorf, Jed McCaleb, and Jake Craige to look at blockchain infrastructure through a developer-tools lens. Konstantin Richter at Blockdaemon moderated this panel at Heavybit, and I sat with Brian Behlendorf (Executive Director, Hyperledger), Jed McCaleb (co-founder, Stellar), and Jake Craige (Lead Developer, Coinbase). I brought the developer-tools-investor lens to the conversation: how does this perform in production, how do you monitor it, what does the surrounding tooling need to mature. We covered the real tradeoffs of decentralization, node management, metrics and monitoring requirements, standardization of distributed systems, and the regulatory picture. The questions I kept returning to were practical ones for the developers building on blockchain infrastructure, the same operational and reliability concerns I had been applying since Amazon and through the DevOps movement. ### Incident Management for Operations (foreword by Jesse Robbins) - O'Reilly Media · Other · 2017-07-01 - Original: https://www.oreilly.com/library/view/incident-management-for/9781491917619/ - Canonical: https://jesserobbins.com/mentions/incident-management-for-operations-schnepp-vidal-hawley-oreilly/ > "This groundbreaking book is the foundation to building an effective operations culture for organizations of any size, with systems of any complexity, and failures of any severity." > — Jesse Robbins, from the foreword I wrote the foreword to Schnepp, Vidal, and Hawley's O'Reilly book applying fire-service incident command to IT operations. ## My foreword to the book This book originated from an argument during my first year as Amazon's "Master of Disaster," as I began applying the incident management and operations practices I learned as a firefighter to improve Amazon's overall reliability and resiliency. I vividly remember facing a room full of scowling engineers and managers who were saying, "I get that these ideas work for firefighters, but do you really think they can work at internet speed?" The answer, of course, was yes. The systems and best practices developed over decades of managing complex emergency incidents, where seconds count and lives are on the line, work just as well for managing complex incidents for technology organizations. Over the next few years, my team and I used these techniques and systems to help transform the culture and technology of what is now one of the greatest engineering and operations organizations in the world. When I left Amazon, it was clear to me that as the world was becoming increasingly connected and distributed, people would come to depend on the new technology we build and systems we run as part of their daily lives. It was also clear to me and a group of passionate peers that there were too few people with the knowledge and experience to build and run these systems at scale. My friend, Artur Bergman helped me found the O'Reilly Velocity Performance & Operations conference to organize, develop, and spread our emerging and critical professional discipline. As Velocity grew, I started sharing my work with friends and mentors in the Fire Service. I am fortunate to have worked with and been trained by some of the most experienced and respected incident management experts in the world, and I asked them to help build and expand on what I had started. I convened the first "Web Ops/Fire Ops" summit on a beautiful day at Artur's loft in San Francisco. Attending from "the internet" were Artur Bergman (Fastly/Wikia), John Adams (Twitter), Johnathan Heiliger (Facebook), Pedro Canahuati (Facebook), Simon Wistow (Fastly), and Christopher Brown (Amazon/Chef/Microsoft). Attending from the "Fire Ops" side were the authors of this book: Chris Hawley, Rob Schnepp, and Ron Vidal. After a few hours of sharing backgrounds, "war stories," and a lot of laughter, it became clear to everyone that there was both the need and opportunity for a tech-oriented incident management training program. Chris, Rob, and Ron formed Blackrock Partners and began consulting with large companies on how to improve their operations. Since then they have worked with dozens of tech companies, trained thousands of new responders, and reviewed hundreds of incidents as they help companies "work like a fire department at internet speed." This book is the first publicly released product of their exceptional work, and is the essential foundation for building technology and organizations that people can depend on. I hope you use it. As we say in the fire department, "See you at the big one!" — Jesse Robbins Founder and CEO, Orion Labs, Inc. ### 'Star Trek Communicator Startup' Sets Out to Build a World Powered by Voice - Wired · Article · 2015-01-21 - Original: https://www.wired.com/2015/01/star-trek-communicator-startup-sets-build-world-powered-voice/ - Canonical: https://jesserobbins.com/mentions/star-trek-communicator-startup-world-powered-by-voice-wired/ Wired covered OnBeep's rebrand to Orion Labs, framing the wearable push-to-talk Onyx as a real-world Star Trek communicator and Jesse Robbins' bet on a world powered by voice for teams that work away from screens. Wired covered our rebrand from OnBeep to Orion Labs and the idea behind it: a world powered by voice. The Star Trek communicator comparison followed Onyx from the day we launched it, and we stopped fighting it. Tap the badge on your chest, talk to your team, anywhere. The product was a wearable, but the company was about voice as an interface for people who work with their hands and eyes busy. Firefighters, nurses, field crews, warehouse teams. That conviction came straight from my time in emergency services, and it is why I cofounded the company. ### This Startup Thinks Your Workplace Needs Wearable Walkie-Talkies - Wired · Article · 2014-11-05 - Original: https://www.wired.com/2014/11/byod-wearables/ - Canonical: https://jesserobbins.com/mentions/jesse-robbins-onbeep-onyx-wired/ Wired profiled Onyx, the $99 wearable walkie-talkie from Jesse Robbins' startup OnBeep. The device clips to clothing, pairs with a smartphone, and gives workplace teams instant push-to-talk voice over Wi-Fi or cellular. Wired covered the launch of Onyx, the wearable walkie-talkie we built at OnBeep. Onyx was a $99 device that clipped to your clothing, paired with your phone, and gave a team instant push-to-talk voice over Wi-Fi or cellular, at any distance. The idea came from my years as a volunteer firefighter. On an incident, voice is the coordination layer: one button, instant connection to the whole team, eyes up the entire time. Most workplaces ran on radios that had not changed in decades, or on phones that bury urgent communication under apps. We wanted to give frontline teams the immediacy of a radio with the reach of the internet. OnBeep became Orion Labs a few months later, and the work continued from there. ### Building Companies that Devs & DevOps Teams Love, And Avoiding Expensive Mistakes - Heavybit · Article · 2013-06-25 - Original: https://www.heavybit.com/library/video/building-companies-that-devs-and-devops-teams-love-and-avoiding-expensive-mistakes - Canonical: https://jesserobbins.com/mentions/building-companies-devops-teams-love-jesse-robbins/ Heavybit talk on the expensive mistakes developer-tools founders make, covering positioning, developer experience, pricing, and go-to-market traps. A 2013 Heavybit talk covering the expensive mistakes developer-tools founders make across positioning, developer experience, pricing, and go-to-market. The lessons cover product positioning, developer experience, pricing, and the go-to-market traps that catch technical founders. I made most of these mistakes firsthand building Chef from an open-source project into an enterprise infrastructure company. Learn a dollar of lesson for every one you spend in failure. ### Tim O'Reilly on Why We Started the Velocity Conference - O'Reilly Radar · Article · 2013-06-17 - Original: http://web.archive.org/web/20221127165413/http://radar.oreilly.com/2013/06/why-we-started-the-velocity-conference.html - Canonical: https://jesserobbins.com/mentions/tim-oreilly-on-why-we-started-velocity-conference/ Tim O'Reilly's 2013 retrospective on how the Velocity Conference began.

A note from Jesse

Tim O'Reilly's 2013 O'Reilly Radar post recalling how the Velocity Conference began. The original URL redirects, so the full text is preserved here.
Back in 2006, Debra Chrapaty, then VP of Operations for Windows Live (later CIO at Zynga, and now CEO of Nirvanix) made a prescient comment to me: "In the future, being a developer on someone's platform will mean being hosted on their infrastructure." As it often turns out, things don't work out quite as planned. A few months later, Amazon announced EC2, and it was Amazon, not Microsoft, that became the platform whose infrastructure startups chose to host their applications on. But Debra certainly nailed the big idea! I wrote a blog post about that conversation, entitled Operations: The New Secret Sauce, which included the statement "Operations used to be thought of as boring. It's now ground zero in the computing wars." Jesse Robbins, then "Master of Disaster" at Amazon and later co-founder and CEO of Opscode, told me that everyone in operations at Amazon printed out that blog post and posted it in their cubicles. Operations had been a relatively low-status job. Jesse told me that was the first time anyone had made a strong public statement about how important it was becoming. As a result of that post, Jesse, Steve Souders, and a group of others came to me the following year and said "We need a gathering place for our tribe." That gathering place became the Velocity Conference, now in its sixth year. We chose to include not just web operations, but also web performance and the emerging field of "DevOps," the development model for applications hosted in the cloud. This seems to be part of the secret sauce of some of our most successful events: the recognition that it's not just about technology but the people who put it into practice. At the heart of conferences like Velocity and Strata are new job descriptions, new skills, and new opportunities to grow careers and companies. That's also why we increasingly think of these events not as conferences but as gathering places for communities. Technology matters. The people who put it into practice matter more. ### Jesse Robbins on the Rise of DevOps (InfoQ Interview) - InfoQ · Article · 2013-01-17 - Original: https://www.infoq.com/interviews/Awesome-DevOps-Jesse-Robbins/ - Canonical: https://jesserobbins.com/mentions/rise-of-devops-jesse-robbins-infoq/ InfoQ interviewed me on how DevOps started, why infrastructure as code changed operations, and what it actually takes to get developers and ops working together. Harry Brumleve interviewed me for InfoQ in January 2013, the point where DevOps was crossing from a niche conversation into a category the broader industry was paying attention to. We covered the transition from traditional operations into DevOps, what infrastructure as code actually meant in practice, and how the operating-model work I had done at Amazon could be made replicable for everyone else. ### DevOps is a reorientation DevOps is a reorientation of how organizations build and run software. The walls between development teams writing code and operations teams running it are failure modes: points where accountability breaks down and blame accumulates instead of improvement. At Amazon, we treated operations as competitive advantage instead of cost center. The question I kept getting in 2013 was how to make that culture replicable beyond the hyperscalers. ### Infrastructure as code By 2013 Chef was the clearest example of what infrastructure as code meant in practice. Servers are instances of a declared state that code creates, modifies, and destroys. The same version control, testing, and review workflows software teams use for applications now apply to infrastructure. A change to a server config becomes a pull request, audited and reproducible, instead of an undocumented action by a sysadmin at 2am. Teams deploy infrastructure the way they deploy code, with confidence the outcome will match the spec. ### Velocity and the practitioners Velocity, which I cofounded with Steve Souders at O'Reilly in 2007, was the room where fierce competitors, Google, Amazon, Microsoft, Facebook, shared what they knew about running complex systems reliably. The practitioners there developed shared vocabulary, shared standards, and shared intuitions about what good operations looked like. By 2013 that work had a name. ### Q&A: Ex-Amazon 'Master of Disaster' Jesse Robbins on the Power of 'Relentless Optimism' in Startups - GeekWire · Article · 2012-10-27 - Original: https://www.geekwire.com/2012/qa-examazon-master-disaster-jesse-robbins/ - Canonical: https://jesserobbins.com/mentions/master-of-disaster-relentless-optimism-geekwire/ > "When you're trying to change the way big organizations work, a lot of people say no a lot. Rather than try to fight them, you've got to find a way to make them say yes. Being a force for awesome in the world is finding ways to say yes." > — Jesse Robbins GeekWire ran a long Q&A while I was running Opscode and pulled out the operating principle I kept using inside Amazon: when people say no, find a way to make them say yes. Jeff Dickey caught me in 2012, mid-stride at Opscode. I had been a technology builder since high school, stepped away from tech to train as a firefighter and EMT, and walked back in through the door at Amazon on August 20, 2001. The Q&A traces that path and lands on the operating principle I had been using inside Amazon and was now using to build Opscode. On 9/11, I woke up in a hospital after emergency surgery and watched the day unfold on a television. That was when I understood that the operational skills I had been training in the fire service translated directly to a technology organization that thousands of people depended on. As I told Jeff: "I decided I'm going to figure out a way to mix these two worlds together." That decision became Master of Disaster, GameDay, and eventually Chef. On driving change in large organizations: "When you're trying to change the way big organizations work, a lot of people say no a lot. Rather than try to fight them, you've got to find a way to make them say yes. Being a force for awesome in the world is finding ways to say yes." On startup life, my advice in 2012 was the same advice I give founders now: "If you're struggling, recognize it's going to be this way forever." The work does not get easier; you get better at it. ### Resilience Engineering: Learning to Embrace Failure - ACM Queue · Article · 2012-09-12 - Original: https://queue.acm.org/detail.cfm?id=2371297 - Canonical: https://jesserobbins.com/mentions/resilience-engineering-learning-embrace-failure-acm-queue/ > "You can't choose whether or not you're going to have failures — they are going to happen no matter what — but you can choose in many cases when you're going to learn the lessons." > — Jesse Robbins Jesse Robbins (Amazon), Kripa Krishnan (Google), and John Allspaw (Etsy) discuss how they built organizations that deliberately trigger failure to get stronger: powering off data centers, running 96-hour disaster simulations, and transforming blame cultures into learning cultures. Three teams. Three companies. Same answer arrived at independently. That is the part of this article that still matters. Tom Limoncelli moderated the discussion and wrote it. Tom Limoncelli is extraordinarily accomplished. His books shaped our profession as it evolved. He pulled three of us into the same room: Kripa Krishnan from Google, John Allspaw from Etsy, and me from Amazon. None of us had read each other first. We had built versions of the same discipline because the problem was the same. Kripa Krishnan ran Google's program, which they called DiRT. By 2012 she had been doing it for about six years. Her exercises ran 72 to 96 hours, hundreds of engineers around the clock, war rooms staffed by about fifty rotating volunteers. The details I have never forgotten are the failures that surfaced. Google brought down a network in São Paulo and watched the links die in Mexico, because nobody knew the dependency was there. A data center where the machines refused to come back online because they had run out of DHCP leases. Kripa had Ben Treynor as her executive sponsor. Ben went on to create a similar program at Google called Site Reliability Engineering (SRE). Without that air cover, none of this happens. John Allspaw was running technical operations at Etsy after stops at Salon.com, Friendster, and Flickr, where he was engineering manager. John brought the academic frame to the conversation. He was the one who put Erik Hollnagel's four cornerstones of resilience on the table: anticipation, monitoring, response, learning. He was the one who said the thing the industry took years to absorb. He had announced publicly that he would not fire an engineer for taking down a site he was responsible for. He described the substitution test Etsy used in postmortems. Pull in an uninvolved engineer, give them the same context, ask what they would have done. Almost every time, the answer is the same command. The problem is not the person. The Brooklyn Bridge line is his too. You do not shut down the whole bridge because one lane is out. What I described was GameDay at Amazon, which I had started in 2003 and 2004, when horizontal scalability across unreliable hardware was not yet a settled idea. We were feeling our way. The exercises were not simulated. We powered facilities off without notice and let the systems fail naturally. In one of them I used my fire service training to script a simulated fire down to the minute, with operators posing as facilities staff calling operations with updates. I had executive support I never took for granted. Werner Vogels was my exec sponsor. Jeff Bezos thought the idea was interesting. The Amazon ops team was the village that actually did the work. Tim O'Reilly later gave this community a place to find each other in public, the Velocity Conference. I cofounded Velocity with Steve Souders after I left Amazon, and later passed the torch to John Allspaw. The discipline framework did not come from tech. I trained as a volunteer firefighter with the Seattle Fire Department before I stepped away from tech. Fire service teaches you that you do not get good at incident response by having an opinion about it. You get good by drilling, and then drilling more, and then drilling again under conditions you do not control. Build the muscle continuously, because the only way to find the failures hiding inside a complex system is to trigger them on purpose, on your terms, before they trigger themselves on theirs. That is the line I gave Tom that has traveled the furthest. You do not get to choose whether you have failures. You get to choose, in many cases, when you learn the lessons. The convergence is what still holds up. Three organizations, no shared playbook, all arriving at the same answer. Stop trying to prevent failure. Build the muscle to absorb it. Never blame the human at the end of the chain for the system around them. What happened after this article ran is the next part of the story. ### Jesse Robbins on DevOps as Business Alignment - Thoughtworks · Video · 2012-08-17 - Original: https://www.thoughtworks.com/insights/blog/jesse-robbins-discusses-devops-and-cloud-computing - Canonical: https://jesserobbins.com/mentions/jesse-robbins-devops-cloud-computing-thoughtworks/ > "The role of operations is the role of enabling as much awesome as you can." > — Jesse Robbins Jez Humble interviewed me at Thoughtworks on DevOps as business alignment: developers, operations, and the company shipping faster without giving up reliability. Jez Humble interviewed me for Thoughtworks in 2012 as part of his continuous delivery video series. The conversation lays out how I was framing DevOps then: business alignment first, collaboration second. ### DevOps as business alignment The starting point is the incentive problem. Developers create value when code reaches production. Operations teams were traditionally rewarded for keeping production unchanged. Most companies treated both groups as support functions instead of seeing delivery and operations as part of the same value stream. DevOps is the shift where operations becomes a way to help the business move faster without giving up reliability. > "Code that is written and not deployed is worthless." ### Operations as enablement I walked Jez through my own pivot from gatekeeper to platform builder. At Amazon, I blocked a launch because I was worried about availability. Neil Roseman overrode me. The site went down for a couple of days. The stock and order rates went up. That was when I realized the better role for operations was to make launches safer and more repeatable, not harder to do. > "The role of operations is the role of enabling as much awesome as you can." ### Chef and infrastructure as code Chef sits inside the larger move toward programmable infrastructure. EC2, Rackspace, OpenStack, VMware, and private-cloud APIs made it possible to provision infrastructure through software. Chef was the glue that let teams describe the desired state of infrastructure and applications together. Infrastructure as code is the moment infrastructure becomes part of the application: configurable, repeatable, testable, close enough to the product that teams can move without waiting on manual provisioning. ### EC2 and the cloud operating model I told Jez the story of EC2's origins inside Amazon, including my own first reaction: I tried to block it. Chris Pinkham and Christopher Brown built the project away from Seattle, in Cape Town, and I gave them a hard time about exposing what I considered my operational perimeter to the public internet. EC2 changed who could ask for infrastructure, how fast they could get it, and what operations teams had to become. ### Simple services, fast feedback The cloud architecture advice from the back half still reads cleanly. Build small services. Keep APIs simple. Push complexity up the stack. Avoid monoliths that force every scaling problem into one place. Prefer rough consensus with running code over elaborate first designs. If the system is split into clear services, the teams can own, operate, and improve those services. Architecture is a way of making responsibility visible. Transcript: https://jesserobbins.com/mentions/jesse-robbins-devops-cloud-computing-thoughtworks.md ### Changing Culture & Being a Force for Awesome - O'Reilly Velocity Conference · Video · 2012-06-28 - Original: https://www.youtube.com/watch?v=OU8ihx3nT6I - Canonical: https://jesserobbins.com/mentions/velocity-2012-changing-culture-force-awesome-oreilly/ > "Don't fight stupid. Focus on where you can make more awesome." > — Jesse Robbins My 2012 Velocity talk on changing engineering culture from the inside. Start small, build champions, use metrics to create confidence, exploit compelling events. This is my Velocity 2012 talk. I had been giving versions of this to smaller rooms since 2011 and earlier at Amazon. Velocity was the room where I tried to say it cleanly to the people who were going to take it back to their own organizations. ### The framework Five steps, each building on the last. 1. **Start small.** Pick the smallest project with receptive people. Call it an experiment. Do not trigger the organizational immune system. 2. **Create champions.** Get your boss on board first. Then spread credit as widely as you can. Let other people feel ownership of the change. 3. **Use metrics to build confidence.** Find one number that supports your change, time from commit to deploy, cost of an outage, and use it ruthlessly to build the case. 4. **Celebrate successes.** Tell the story with data. Be positive about people. Leave room for resistors to come around without losing face. 5. **Exploit compelling events.** When the site goes down or a compliance mandate lands, use the moment to push for the change you have been building toward. ### The Katrina lesson on permission I tell a story from my deployment as a task force leader during Hurricane Katrina. A volunteer kitchen staffed by anarchists was feeding thousands of people a day, and FEMA kept trying to shut them down because no one would say who was in charge. The fix was simple. Make every volunteer a "site director." When FEMA asked who was in charge, someone would answer "I'm a site director," and FEMA would deliver supplies. The lesson stuck with me: most of the time when people are saying no, what they really mean is they do not know how to say yes. I used the same move at Amazon by typing "Master of Disaster" into a form as my job title. It stuck. ### The rule The talk keeps returning to one line: don't fight stupid. Focus on where you can make more awesome. Transcript: https://jesserobbins.com/mentions/velocity-2012-changing-culture-force-awesome-oreilly.md ### Jesse Robbins on the State of Infrastructure Automation - O'Reilly Radar · Article · 2012-05-11 - Original: https://web.archive.org/web/2013/http://radar.oreilly.com/2012/05/jesse-robbins-infrastructure-automation.html - Canonical: https://jesserobbins.com/mentions/jesse-robbins-infrastructure-automation-oreilly-radar/ O'Reilly Radar interviewed me on Chef's evolution from open-source project to enterprise infrastructure automation, and where cloud operations was headed next. Tim O'Brien interviewed me for O'Reilly Radar in May 2012, while I was Chief Community Officer at Opscode. We talked about Chef's evolution from a tool the early adopters loved into something enterprise customers were buying at scale, and about the broader shift toward treating infrastructure as code. ### 5 Pivotal Documents in the Evolution of the DevOps Movement - SiliconANGLE · Article · 2012-04-03 - Original: https://siliconangle.com/2012/04/03/5-pivotal-documents-in-the-evolution-of-the-devops-movement/ - Canonical: https://jesserobbins.com/mentions/five-pivotal-documents-devops-movement-devopsangle/ > "Operations is a Competitive Advantage (Secret Sauce for Startups!)" > — Jesse Robbins (cited document) Klint Finley's 2012 canon of DevOps. The Agile Manifesto, Tim O'Reilly's Operations piece, my 'Operations is a Competitive Advantage' post from 2007, John Allspaw's 10 Deploys talk, and Jay Lyman's analyst report. ### The Convergence of DevOps - IT Revolution · Article · 2012-03-30 - Original: https://itrevolution.com/articles/the-convergence-of-devops/ - Canonical: https://jesserobbins.com/mentions/convergence-of-devops-itrevolution/ John Willis's history of how DevOps came together: Agile Infrastructure, Velocity, and Lean Startup as the three threads that converged. John traces three threads that ran in parallel in the late 2000s and converged into the DevOps movement: Agile Infrastructure, Velocity, and Lean Startup. [Patrick Debois](https://jedi.be), Andrew Clay Shafer, and the agile system administration community on one strand. The Velocity Conference at O'Reilly on another. Eric Ries and the Lean Startup community on a third. The Velocity strand at O'Reilly grew out of a 2007 Radar post I wrote applying the technical-debt frame to operations and a Tim O'Reilly follow-up titled "Operations: The New Secret Sauce." DevOpsDays Mountain View 2010, organized by Patrick, John Willis, Andrew Clay Shafer, and Damon Edwards, was the first US DevOpsDays. I was honored to be there. ## Further Reading - [Operations Is a Competitive Advantage](/mentions/operations-competitive-advantage-oreilly-radar/) — The 2007 O'Reilly Radar post in the Velocity thread - [The Art of Web Operations](/mentions/velocity-art-of-web-operations-oreilly-radar/) — Jesse Robbins and John Allspaw at Velocity 2009 - [GameDay: Creating Resiliency Through Destruction](/mentions/gameday-creating-resiliency-through-destruction-usenix/) — USENIX LISA'11 - [5 Pivotal Documents in the Evolution of the DevOps Movement](/mentions/five-pivotal-documents-devops-movement-devopsangle/) — DevOpsANGLE's analysis ### GameDay: Creating Resiliency Through Destruction - USENIX · Talk · 2011-12-20 - Original: https://www.youtube.com/watch?v=zoz0ZjfrQ9s - Canonical: https://jesserobbins.com/mentions/gameday-creating-resiliency-through-destruction-usenix/ > "You don't choose the moment, the moment chooses you. You only choose how prepared you are when it does." > — Jesse Robbins My USENIX LISA'11 talk on GameDay: deliberately inject failures into production to build organizational resilience before real outages happen. I had been running these exercises at Amazon since 2003. My USENIX LISA 2011 talk on GameDay. Boston, December 2011. I opened with the story my fire chief told me on my first day of firefighter academy. He said when you go home tonight and tell your neighbor you are becoming a firefighter, something changes. At 2:00 in the morning when their kid starts choking, they will not dial 911. They will pound on your door with a slumped over kid, looking to you to do something. "Welcome to the other side of 911." And then he said the thing that set the arc of my entire career: "You don't choose the moment. The moment chooses you. You only choose how prepared you are when it does." That is operations. We are the ones they call. We are the ones people look to when things are broken. The core argument of the talk is simple. Resilience is a property of your entire system, and the system includes people, culture, processes, applications, infrastructure, and hardware. People and culture are the most important part. That is weird for a room full of us who tend to shy away from human interaction, but it is the truth. GameDay is about changing people to be resilient. I created GameDay at Amazon after joining in 2001 while testing for the Seattle Fire Department. My title was "Master of Disaster." I owned website availability for every property that bore the Amazon name. I realized the way we were running operations was not going to scale, and I began adapting fire service incident management processes, training, drilling, and fire prevention concepts to make Amazon operate like a fire department. I substituted the word "management" for "command" and created an incident management system that was a word-for-word adaptation of the fire department's incident command system. Werner Vogels was my internal executive sponsor. The methodology has three stages. First, preparation. You identify and mitigate risks. You walk through the system and find the stupid stuff. The oil barrels next to the core database servers. The single points of failure nobody talks about. This alone reduces failure frequency and improves recovery time. Second, you run the drill. This is where most organizations fail. They do a lightweight tabletop exercise, declare victory, and never go further. If you never go all the way to a full-scale live failure exercise, you do not get the confidence that comes from responding to a stressful situation at full speed. As a firefighter, I have been through countless live fire drills. The fire will kill you. People die in training. I can only be an effective firefighter having fought fire. You also discover that systems and processes you thought would work do not. Cell phone systems have single points of failure. Nobody has a printout of everyone's phone number. The most basic things only surface under real stress. Third, you expose latent defects. These are the impossible failures. The ones that cannot happen until they do. They sit underneath the waterline of your systems, and you cannot discover them any other way. They are immediately recognizable in hindsight. "Oh, we totally should have known there was a dependency on that developer's desktop." You do not get to choose whether you have latent defects. You do get to choose when you discover some of them. Run the exercise off-peak rather than finding out on your most important traffic day. The progression matters. You start small. You work with the smallest group of developers who are receptive, and you break something that is a little scary but not too scary. Think of the kid with the fire hose. These hoses produce 80 to 100 pounds of back pressure. You do not hand a recruit a high-pressure fire line on day one. You give them a garden hose, a small pan fire, and you let them succeed. Then they tell everyone how awesome it was. You build on those successes. Then you move up to the full-scale live fire exercise. You pick the worst survivable scenario. I recommend a full data center power-down. It will terrify everyone. Everyone will hate you during the planning phase. It is going to be good for them. You give them a couple months notice. You tell them the date, the time, the facility. They have months to remediate. And then the week comes, and people ask if you are really going to power it down. Yes. You are really going to power it down. You can slip a date, but never cancel. Otherwise no one will ever believe you again. I always power it down. The first time will be a disaster. You will learn more from that one exercise than from years of tabletop reviews. Probably you will learn that your database masters do not come back up after an EPO reset. You will be glad you chose when to learn that. The reason all of this works is the OODA loop, the observe-orient-decide-act cycle that John Boyd described from fighter pilot training. In any crisis, you follow a predictable response. You observe what is going on. You orient yourself based on your training, your experience, whether you have been exposed to situations like this before. If you have never been through a full-scale outage, you will lock up. There are predictable failure modes. I can tell you what you will do. The only reason I am different is I have been through it a lot. GameDay became an internal competitive advantage at Amazon. One team would say, "Go ahead, rip it out. We can power stuff off all day." The other teams wanted to get there. That is the confidence-driven currency for change John Allspaw talks about, where people say "We are better because this happened," instead of the crisis-driven kind where a bad outage opens a short window for sweeping fixes. Every large-scale web operation has since learned some version of this or perished. Google adopted it. Engineers who had been at Amazon brought the practice to Netflix and built [Chaos Monkey](https://netflix.github.io/chaosmonkey/), which randomly terminated instances in production. They expanded it into the [Simian Army](https://netflixtechblog.com/the-netflix-simian-army-16e57fbab116) and later [Chaos Kong](https://netflixtechblog.com/chaos-engineering-upgraded-878d341f15fa) for regional failover testing. Facebook, Yahoo, and dozens of others built their own programs. AWS turned GameDay into a core operational practice and eventually a customer-facing program. The [Well-Architected Framework](https://wa.aws.amazon.com/wat.concept.gameday.en.html) now defines "game day" as a formal reliability concept. [AWS Fault Injection Service](https://aws.amazon.com/fis/) is a managed chaos engineering service that automates the kind of fault injection I was doing by hand. The FIS team [runs their own game days](https://aws.amazon.com/blogs/mt/learn-from-aws-fault-injection-service-team-approach-to-game-days/) using the service they built. That is the right kind of recursion. The practices I described in this talk, controlled fault injection, pre-announced failure exercises, progressive escalation, and blameless post-incident review, became core patterns in what the industry later formalized as [chaos engineering](https://en.wikipedia.org/wiki/Chaos_engineering). Transcript: https://jesserobbins.com/mentions/gameday-creating-resiliency-through-destruction-usenix.md ### Meet 2011 TR35 Winner Jesse Robbins - MIT Technology Review · Video · 2011-12-02 - Original: https://www.youtube.com/watch?v=s55G8eDHGgY - Canonical: https://jesserobbins.com/mentions/meet-2011-tr35-winner-jesse-robbins-mit-tr/ MIT Technology Review interviewed me as a 2011 TR35 honoree, recognizing the work on web operations, infrastructure automation, and reliability at Opscode. ## From the MIT Technology Review TR35 honoree page **Category:** Internet & web **Year Honored:** 2011 **Organization:** Opscode **Region:** Global **Focus:** Fault-tolerant online infrastructure ### Biography Jesse Robbins applied for two jobs in 2001: a Seattle bus driver position and a backup systems engineer role at Amazon.com. Amazon's offer came first, beginning a decade of work on how web companies operate complex server and software networks at scale. Drawing from his background as a volunteer firefighter, Robbins brought crisis management principles to infrastructure design. He recognized that massive global operations inevitably experience failures and built systems to withstand them safely. Rather than preventing failures, he made Amazon resilient to them through architectural fault tolerance and live operational drills that tested teams by temporarily taking entire data centers offline, without affecting customer experience. After leaving Amazon in 2006, Robbins shared his methodologies through blogging. In 2007, he cofounded Velocity, now an annual conference where major competitors openly discuss infrastructure management. Robbins cofounded Opscode in 2008. The company's flagship product, Chef, is an open-source framework for cloud-based infrastructure automation. One notable application involved scientists using Chef to deploy a 10,000-processor supercomputing cluster in 45 minutes on Amazon's cloud, completing complex protein-binding research in eight hours, then shutting down operations, all at a fraction of traditional supercomputing costs. Transcript: https://jesserobbins.com/mentions/meet-2011-tr35-winner-jesse-robbins-mit-tr.md ### The Chef, the Puppet, and the Sexy IT Admin - Wired · Article · 2011-10-26 - Original: https://www.wired.com/wiredenterprise/2011/10/chef_and_puppet/ - Canonical: https://jesserobbins.com/mentions/chef-and-puppet-wired-enterprise/ Wired Enterprise covered the rivalry between Chef and Puppet as infrastructure automation went mainstream, placing Jesse Robbins and Opscode at the center of the industry's shift to infrastructure as code. Cade Metz wrote about the rivalry between Chef and Puppet for Wired Enterprise in 2011. Mainstream tech press covering configuration management was new. For years the only people who cared about infrastructure automation were the ones carrying pagers. Adam Jacob, Barry Steinglass, Nathan Haneysmith, and I cofounded Chef to bring the automation Google and Amazon kept as closely guarded secrets to everyone else. Chef let engineering teams define infrastructure as code, writing recipes to configure entire fleets of servers instead of managing them by hand. Puppet was working the same problem with a different philosophy, and the comparison became a running storyline as the category grew. Cade Metz's Wired Enterprise piece compared Chef and Puppet as configuration management moved from an insider practice to a mainstream category. ### DevOps Cafe Episode 19: Jesse Robbins - DevOps Cafe Podcast · Podcast · 2011-09-20 - Original: http://devopscafe.org/show/2011/9/20/devops-cafe-episode-19.html - Canonical: https://jesserobbins.com/mentions/devops-cafe-episode-19-jesse-robbins/ Damon Edwards and John Willis hosted me on DevOps Cafe to walk through the path from teenage ISP work to firefighting to Amazon to Chef and Velocity. Damon Edwards and John Willis hosted me on DevOps Cafe for a long conversation: teenage ISP work, the Seattle Fire Department, joining Amazon, building GameDay, cofounding Velocity, and starting Chef. A full transcript ran on the Chef blog at the time. ### Puppet, Chef Ease Transition to Cloud Computing - BusinessWeek · Article · 2011-09-01 - Original: https://web.archive.org/web/20110924131145/http://www.businessweek.com/magazine/puppet-chef-ease-transition-to-cloud-computing-09012011.html - Canonical: https://jesserobbins.com/mentions/puppet-chef-ease-transition-cloud-businessweek/ > "The custom tools built by Google, Amazon, and some other guys were such closely guarded secrets. Our founding thesis was to open up these tools to everyone else." > — Jesse Robbins Olga Kharif and Ashlee Vance on Puppet and Chef as the open-source tools bringing infrastructure automation to the enterprise cloud market.

A note from Jesse

Olga Kharif and Ashlee Vance wrote this for [BusinessWeek](https://web.archive.org/web/20110924131145/http://www.businessweek.com/magazine/puppet-chef-ease-transition-to-cloud-computing-09012011.html) on September 1, 2011. The original BusinessWeek URL is gone. The full piece is preserved here from archive.org because it captures the moment infrastructure automation crossed into the enterprise business press, with Opscode's founding thesis on the record.
**Puppet, Chef Ease Transition to Cloud Computing** *By Olga Kharif and Ashlee Vance — BusinessWeek, September 1, 2011* Organizations as diverse as Northrop Grumman, Harvard University, Zynga, and the New York Stock Exchange have filled job websites with requests for talented puppeteers and master chefs. A quick dig into the job listings reveals that these positions have nothing to do with office entertainment or gourmet meals. Instead, the companies want people who have mastered Puppet or Chef, competing software tools that sit at the heart of the cloud computing revolution. In essence, Puppet and Chef are levers used to control data center computers in a more automated fashion. The software has helped companies tap vast stores of computing power in new ways, accelerating research in fields such as financial modeling and genetics. "This really changes the way science gets done," says Jason Stowe, the chief executive officer of Cycle Computing, a startup that uses Chef to configure thousands of computers at a time so that clients can perform calculations at supercomputer speeds. Before adopting Chef, doing such configurations took hours or even days. "We're down to single-digit minutes now," Stowe says. The need for such tools originated with Google, Amazon.com, and their peers, who have long had to deal with the burden of managing tens or even hundreds of thousands of servers to support vast Web operations. Over the years these companies developed custom tools that can quickly turn, say, a thousand new servers into machines capable of displaying Web pages or handling sales. These programs allow the companies to run enormous, $500 million computing centers with about three dozen people at each one. As more and more businesses move their software applications to the cloud, a handful of startups have developed mainstream versions on such data-center software. Puppet and Chef are the two with the highest profile. "The custom tools built by Google, Amazon, and some other guys were such closely guarded secrets," says Jesse Robbins, co-founder of Opscode, the 20-person, Seattle-area startup behind Chef. The company has raised $13.5 million in venture capital. "Our founding thesis was to open up these tools to everyone else." Opscode's Chef and its competitor, built by Puppet Labs, are both open source: Anyone is free to use and adapt the software. The companies make money by selling polished versions of the core technology and additional features, and by charging for advice on how to implement and best use it. Luke Kanies came up with the idea for Puppet in 2003 after getting fed up with existing server-management software in his career as a systems administrator. In 2005 he quit his job at BladeLogic and spent the next 10 months writing code to automate the dozens of steps required to set up a server with the right software, storage space, and network configurations. He formed Puppet Labs to begin consulting for some of the thousands of companies using the software, the list includes Google, Zynga, and Twitter, and earlier this year he released the first commercial version. Stanford University used to rely on a hodgepodge of tools to manage its hundreds of servers. They've since replaced that unorganized toolbox with Puppet. "We were a ragtag team, and now we are a cohesive unit, and our servers require a lot less attention," says Digant Kasundra, an infrastructure systems software developer at the university. Palo Alto-based Jive Software has used Puppet to double the number of servers a single engineer can handle. "It's a huge impact for us," says Matt Tucker, chief technology officer and co-founder of the company. Rivalry between Chef and Puppet is fierce. Puppet Labs argues that its software requires less training and collects more data about what's happening on the network. Chef claims a bigger developer base. Michael Dell, the founder and CEO of Dell, follows Kanies on Twitter. Traditional data-center heavyweights such as Hewlett-Packard and IBM have shown interest in this type of software and could emerge as potential acquirers. The bottom line: Investors have bet $20.5 million that Puppet and Chef, competing server-management tools, will be at the forefront of cloud computing. ### DevOps Culture Hacks: Infecting your Boss & your Business with Awesome - DevOpsDays · Talk · 2011-03-08 - Original: https://legacy.devopsdays.org/events/2011-boston/proposals/devops%20culture%20hacks/ - Canonical: https://jesserobbins.com/mentions/devops-culture-hacks-devopsdays-boston/ > "Don't fight stupid, make more awesome." > — Jesse Robbins DevOpsDays Boston 2011. I gave the culture hacks talk for the first time, no slides, no video, just the framework I had figured out the hard way at Amazon. DevOpsDays Boston, March 2011. The culture hacks framework, presented without slides and without video. We did not call it DevOps at the time. We did not have a word for it. I had been the stereotypical evil, nasty, mean ops guy. I took every outage personally. I had a record of 167 pages for heightened severity incidents in a single 24-hour period. I had multiple stretches of working 72 hours or more recovering from outages. People were afraid of me. I was proud of that, which tells you something about where my head was. The story that made the room understand the problem was about Neil Roseman. I called Neil the VP of Awesome. Every cool project at Amazon seemed to be under Neil. Kindle. Search Inside the Book. A bunch of others. We had a battle over a deploy that I knew would take the site down. I did what every good ops person would do. I said absolutely not. Neil overrode me. He said, "The website may go down, but the stock price will go up." The site went live, and a few seconds later it crashed. Two days of chaos. And the stock and order rates went up. At year end, I was penalized for the outage and for getting in the way of development. The dev teams were rewarded for shipping. That is the fundamental disconnect. Penalized for something out of my control. Rewarded for deploying and creating value. That misalignment of incentives is the root cause of most of the stupid things organizations do. I became so famous for saying "No" that I would sign the Amazon launch posters with a big "No" and a scribble. It was the only value I could create. That was not progress. So I figured out the formula. Five steps, each building on the last. Start small. Find the smallest group of people who are already excited. Call it an experiment. Do not trigger the organizational immune system. I learned this in the fire service. I always tell people I will take 100 percent of the blame for whatever goes wrong as long as we make space to try. Create champions. Get your boss on board first. Give everyone else the credit. You can accomplish anything you want so long as you do not require credit or compensation. At Amazon, I created the Call Leader Program to train senior people to run high-severity incidents. It became a high-status thing. Managers wanted in. Use metrics to build confidence. Find a number that supports your change and use it ruthlessly. Tell the story with data. Have your champions evangelize on your behalf. Celebrate successes. Create moments in time where people recognize that a change has occurred and that change is good. Exploit compelling events. Big outages create cultural permission to make sweeping changes. And sometimes you create the compelling event yourself. That is what GameDay was. I got executive sponsorship, created a program where we broke critical parts of the infrastructure, and suddenly everyone needed something from me. Compelling event, manufactured. I refined this talk at Velocity in 2012, but this was the first time I said it all out loud to a room of practitioners. The pattern has not changed. Don't fight stupid. Make more awesome. ## Further Reading - [Changing Culture and Being a Force for Awesome](/mentions/velocity-2012-changing-culture-force-awesome-oreilly/) — the refined Velocity 2012 version of this talk, with video - [GameDay: Creating Resiliency Through Destruction](/mentions/gameday-creating-resiliency-through-destruction-usenix/) — USENIX, 2011 - [Operations Is a Competitive Advantage](/mentions/operations-competitive-advantage-oreilly-radar/) — the 2007 O'Reilly Radar post that framed operations as strategic - [The Convergence of DevOps](/mentions/convergence-of-devops-itrevolution/) — John Willis traces the threads that created DevOps - [An Oral History of #HugOps](/mentions/oral-history-hugops-protocol/) — Protocol, 2021 Transcript: https://jesserobbins.com/mentions/devops-culture-hacks-devopsdays-boston.md ### MIT Technology Review TR35: Innovators Under 35 - MIT Technology Review · Article · 2011-01-01 - Original: http://www2.technologyreview.com/tr35/profile.aspx?TRID=1108 - Canonical: https://jesserobbins.com/mentions/tr35-jesse-robbins-technology-review/ The MIT Technology Review TR35 listing for 2011, citing my work on web operations, cloud, and resilience engineering at Amazon and Opscode. MIT Technology Review named me to TR35 in 2011, citing the GameDay program at Amazon and the founding work at Opscode. The shorter, video version of the same recognition lives at [Meet 2011 TR35 Winner Jesse Robbins](/mentions/meet-2011-tr35-winner-jesse-robbins-mit-tr/). ### Web Operations: Keeping the Data on Time - O'Reilly Media · Other · 2010-06-28 - Original: https://www.oreilly.com/library/view/web-operations/9781449377465/index.html - Canonical: https://jesserobbins.com/mentions/web-operations-book-allspaw-robbins-oreilly/ > "The Web is changing the way we live and touches every person alive. As more and more people depend on the Web, they depend on us. Web Operations is work that matters." > — Jesse Robbins, from the foreword John Allspaw and I co-edited this O'Reilly book of essays from practitioners at Amazon, Google, and Flickr on running large web sites. John Allspaw and I co-edited *Web Operations: Keeping the Data on Time* and O'Reilly published it in June 2010. It is a book of essays from people who were actually doing this work at the time, written for the people who would do it next. The book grew out of the [Velocity Conference](https://www.oreilly.com/conferences/velocity.html) community I cofounded at O'Reilly in 2008. Velocity gave the people running the largest sites on the web a place to compare notes in public for the first time. Web Operations is what those notes looked like in book form. The thesis was that operating large systems is its own engineering discipline, not a chore tacked onto development. That position was contested at the time. It is now the consensus, and the lineage from this book runs through DevOps, the SRE books from Google, and the platform engineering work the CNCF formalized a decade later.

My foreword to the book

It's been over a decade since the first websites reached real scale. We were there then, in those early days, watching our sites growing faster than anyone had seen before or knew how to manage. It was up to us to figure out how to keep everything running, to make things happen, to get things done. While everyone else was at the launch party, we were deep in the bowels of the datacenter racking and stacking the last servers. Then we sat at our desks late into the night, our faces lit with the glow of logfiles and graphs streaming by. Our experiences were universal. Our software crashed or couldn't scale. The databases crashed and data was corrupted, while every server, disk, and switch failed in ways the manufacturer absolutely, positively said it wouldn't. Hackers attacked, first for fun and then for profit. And just when we got things working again, a new feature would be pushed out, traffic would spike, and everything would break all over again. In the early days, we used what we could find because we had no budget. Then we grew from mismatched, scavenged machines hidden in closets to megawatt-scale datacenters spanning the globe filled with the cheapest machines we could find. As we got to scale, we had to deal with the real world and its many dangers. Our datacenters caught fire, flooded, or were ripped apart by hurricanes. Our power failed. Generators didn't kick in, or started and then ran out of fuel, or were taken down when someone hit the Emergency Power Off. Cooling failed. Sprinklers leaked. Fiber was cut by backhoes and squirrels and strange creatures crawling along the seafloor. Man, machine, and Mother Nature challenged us in every way imaginable and then surprised us in ways we never expected. We worked from the instant our pagers woke us up or when a friend innocently inquired, "is the site down?" or when the CEO called scared and furious. We were always the first ones to know it was down and the last to leave when it was back up again. Always. Every day we got a little smarter, a little wiser, and learned a few more tricks. The scripts we wrote a decade ago have matured into tools and languages of their own, and whole industries have emerged around what we do. The knowledge, experiences, tools, and processes are growing into an art we call Web Operations. We say that Web Operations is an art, not a science, for a reason. There are no standards, certifications, or formal schooling (at least not yet). What we do takes a long time to learn and longer to master, and everyone at every skill level must find his or her own style. There's no "right way," only what works (for now) and a commitment to doing it even better next time. The web is changing the way we live and touches every person alive. As more and more people depend on the web, they depend on us. Web Operations is work that matters. — Jesse Robbins
## The chapters ## The chapters From John Allspaw's preface, "How This Book Is Organized": - **Chapter 1, Web Operations: The Career** by [Theo Schlossnagle](/people/theo-schlossnagle/). What this field actually encompasses, and why the skills needed are gained by experience more than by formal education. - **Chapter 2, How Picnik Uses Cloud Computing: Lessons Learned** by [Justin Huff](/people/justin-huff/). How Picnik.com deployed and sustained its infrastructure on a mix of on-premise hardware and cloud services. - **Chapter 3, Infrastructure and Application Metrics** by [Matt Massie](/people/matt-massie/) and [John Allspaw](/people/john-allspaw/). The importance of gathering metrics from both your application and your infrastructure, and considerations on how to gather them. - **Chapter 4, Continuous Deployment** by [Eric Ries](/people/eric-ries/). The advantages of deploying code to production in small batches, frequently. - **Chapter 5, Infrastructure as Code** by [Adam Jacob](/people/adam-jacob/). An overview of the theory and approaches for configuration and deployment management. - **Chapter 6, Monitoring** by [Patrick Debois](/people/patrick-debois/). The various considerations when designing a monitoring system. - **Chapter 7, How Complex Systems Fail** by Dr. [Richard Cook](/people/richard-cook/). His whitepaper on systems failure and the nature of complexity often found in web architectures, with web operations-specific notes added to the original. - **Chapter 8, Community Management and Web Operations**. [John Allspaw](/people/john-allspaw/)'s interview with [Heather Champ](/people/heather-champ/) on how outages and degradations should be handled on the human side of things. - **Chapter 9, Dealing with Unexpected Traffic Spikes** by [Brian Moon](/people/brian-moon/). Experiences with huge traffic deluges at Dealnews.com and what they did to mitigate disaster. - **Chapter 10, Dev and Ops Collaboration and Cooperation** by [Paul Hammond](/people/paul-hammond/). Places where development and operations can come together to enable the business, both technically and culturally. - **Chapter 11, How Your Visitors Feel: User-Facing Metrics** by [Alistair Croll](/people/alistair-croll/) and [Sean Power](/people/sean-power/). Metrics that can be used to illustrate what the real experience of your site is. - **Chapter 12, Relational Database Strategy and Tactics for the Web** by [Baron Schwartz](/people/baron-schwartz/). Common approaches to database architectures and some pitfalls that come with increasing scale. - **Chapter 13, How to Make Failure Beautiful: The Art and Science of Postmortems** by [Jake Loomis](/people/jake-loomis/). What makes or breaks a good postmortem and root cause analysis process. - **Chapter 14, Storage** by [Anoop Nagwani](/people/anoop-nagwani/). The gamut of approaches and considerations when designing and maintaining storage for a growing web application. - **Chapter 15, Nonrelational Databases** by [Eric Florenzano](/people/eric-florenzano/). Considerations and advantages of using a growing number of "nonrelational" database technologies. - **Chapter 16, Agile Infrastructure** by [Andrew Clay Shafer](/people/andrew-clay-shafer/). The human and process sides of operations, and how agile philosophy and methods map (or not) to the operational space. - **Chapter 17, Things That Go Bump in the Night (and How to Sleep Through Them)** by [Mike Christian](/people/mike-christian/). The various levels of availability and Business Continuity Planning (BCP) approaches and dangers. ### Ex-Amazon 'Master of Disaster' Animates Server Chef - The Register · Article · 2010-06-22 - Original: https://www.theregister.com/off-prem/2010/06/22/ex-amazon-master-of-disaster-animates-server-chef/951477 - Canonical: https://jesserobbins.com/mentions/ex-amazon-master-of-disaster-animates-server-chef-register/ The Register profiled my move from Amazon's Master of Disaster role to co-founding Opscode and launching Chef, tracing the line from reliability engineering to infrastructure as code. The Register profiled my move from Amazon to Opscode under the headline "Ex-Amazon 'Master of Disaster' Animates Server Chef." The article introduced the "Master of Disaster" framing to a global technology audience and connected it to Chef, the open-source infrastructure automation framework we had just launched. At Amazon, my title was Master of Disaster. The Register put the job plainly: I was "responsible for website availability for every property bearing the Amazon brand." The practice we built for that work was GameDay: schedule failure on purpose, run the drill, expose the latent defects, do it again. That same orientation shaped Chef. Instead of configuring servers one at a time, engineers wrote "recipes" that programmatically configured and managed fleets of servers across cloud providers like Amazon EC2 and in private data centers. The Register called it "object-oriented programming for system administrators." Adam Jacob and I co-founded Opscode with Nathan Haneysmith and Barry Steinglass, betting that infrastructure could be versioned, tested, and deployed like application code. Chef grew to serve Apple, Facebook, Google, and IBM, among many others, before Progress Software acquired the company in 2020. ### The Origins of Amazon's Cloud Computing - GigaOM · Article · 2010-06-18 - Original: https://web.archive.org/web/2013/http://gigaom.com/2010/06/18/the-origins-of-amazons-cloud-computing/ - Canonical: https://jesserobbins.com/mentions/origins-amazon-cloud-computing-gigaom/ > "I was horrified at the thought of the dirty, public Internet touching MY beautiful operations." > — Jesse Robbins I told Stacey Higginbotham at GigaOM the actual origin of EC2. Chris Pinkham wanted to keep working from South Africa, and I, running ops at Amazon, was at first horrified by the idea. Stacey Higginbotham at GigaOM reconstructed the actual origin of Amazon's cloud computing platform. Chris Pinkham, an Amazon engineer, wanted to return home to South Africa, and Amazon agreed to let him keep working from Cape Town. Pinkham and Christopher Brown built the small remote team that designed and shipped the first virtualized server platform inside Amazon, the system that became EC2. I was running availability at Amazon at the time, and I told Stacey I had initially resisted the project. "I was horrified at the thought of the dirty, public Internet touching MY beautiful operations." Pinkham and Brown built their platform in a separate data center, outside my operational perimeter. Werner Vogels was the executive sponsor who created the space for them to do it. The piece ends with a line from Carl Brooks that the operations posture at Amazon helped create the conditions for the platform that would invert it, an early instance of the pattern I had named three years earlier on O'Reilly Radar: [you become what you disrupt](/about/you-become-what-you-disrupt/). ### Velocity: The Art of Web Operations - O'Reilly Radar · Article · 2009-06-22 - Original: https://web.archive.org/web/2013/http://radar.oreilly.com/2009/06/velocity-the-art-of-web-operat.html - Canonical: https://jesserobbins.com/mentions/velocity-art-of-web-operations-oreilly-radar/ Tim O'Reilly's note opening Velocity 2009 tells the origin story: Steve Souders, Andy Oram, and I asked for a conference for our community. Two years in, 700 people showed up.

A note from Jesse

Tim wrote this the week Velocity 2009 opened. The original is gone from O'Reilly Radar, so I am preserving it. He recounts how Steve Souders, Andy Oram, and I asked for a conference for our community.
Two years ago, at the 2007 [O'Reilly Open Source Convention](http://conferences.oreilly.com/oscon), a group of web operations professionals, led by [Jesse Robbins](http://radar.oreilly.com/jesse/) and [Steve Souders](http://www.oreillynet.com/pub/au/2951) along with O'Reilly editor [Andy Oram](http://www.oreillynet.com/pub/au/36), asked for a meeting with me. Their message: "We need a separate conference for our community." That community: the web operations professionals who keep sites up and running. They knew I was receptive. A year earlier, I'd published a blog post entitled [Operations: The New Secret Sauce](http://radar.oreilly.com/2006/07/operations-the-new-secret-sauc.html). I had been pushing for years to get books on web operations into our publishing list (and in fact, Steve's book, [High Performance Websites](http://oreilly.com/catalog/9780596529307/) was in production at that time, and Andy had a number of other titles in the works.) But nonetheless, the meeting felt like an intervention. It was absurdly exciting. I had been thinking in the abstract about the fact that as we move to a software as a service world, one of the big changes was that applications had people "inside" of them, managing them, tuning them, and helping them respond to constantly changing conditions. The skills and tools used by these people would need to be spread to a wider audience. But here were a group of these people, a big group, saying "We need an identity as a profession, and we need a gathering place for our tribe. We want your help." How could I say no? We agreed to start with a "Summit" meeting to bring together the community and brainstorm ideas. [Gina Blaber](http://twitter.com/ginablaber), our VP of Conferences, organized a meeting of 30 or 40 of the "big dogs", and the excitement was palpable. She moved quickly on from there to launch the [Velocity conference](http://en.oreilly.com/velocity2009). It was a success in its first year out, and the second annual conference, starting today in San Jose, promises to be even better. What's more, the fact that attendance has surpassed last year, in an economy that has depressed attendance at many industry conferences by 30-50%, says something about the growing importance of this new field. Back when I first began thinking about this topic for our publishing program, five or six years ago, observing that we needed books on what went on inside of Google and companies like it, the tools and processes they use to deliver such astounding performance and scale, the pushback was that "there are only a few companies operating at that scale." But of course, if you've heard me speak, you've probably heard me quote William Gibson: ["The future is already here. It's just not evenly distributed yet."](http://en.wikipedia.org/wiki/William_Gibson#cite_note-126) Now, there are hundreds of companies (at least) operating at the scale that Google was operating at when I first made those statements. Over 700 of the people who keep them running are converging on San Jose today. ### Five Whys: Try to Learn a Dollar's Worth of Lesson for Every One You Spend in Failure - Venture Hacks · Quote · 2008-11-17 - Original: https://venturehacks.com/five-whys-2 - Canonical: https://jesserobbins.com/mentions/five-whys-jesse-robbins-quote-venturehacks/ > "Try to learn a dollar's worth of lesson for every one you spend in failure." > — Jesse Robbins Eric Ries quoted me in his Venture Hacks guide to Five Whys: try to learn a dollar's worth of lesson for every dollar spent in failure. The line came from Amazon GameDay practice. Eric Ries quoted me in his November 2008 Venture Hacks post on implementing Toyota's Five Whys methodology at startups. The line he pulled was the failure-as-learning maxim I repeated about Amazon GameDay. The idea behind it is what GameDay was built on. The cost of an outage is fixed once it happens. The learning you extract from it is not. We ran scheduled, deliberate failures at Amazon to force the learning before the unscheduled failures could take it. Eric saw the same shape inside a startup: a failed experiment is not waste if you pull proportional learning out of it, and Five Whys was his mechanism for making that extraction systematic instead of accidental. ### The Do-Good Imperative - BusinessWeek · Article · 2008-07-07 - Original: https://web.archive.org/web/20080712045701/http://www.businessweek.com/technology/content/jul2008/tc2008076_973163.htm - Canonical: https://jesserobbins.com/mentions/do-good-imperative-businessweek/ > "One of the interesting things with being a pretty senior technology person operating in a disaster is that you get to see the state of the art versus the state of the practice." > — Jesse Robbins BusinessWeek's CEO Guide to Disaster Readiness covered my work at the seam between emergency response and technology, including the Velocity conference I co-founded. BusinessWeek built a special report in July 2008 called the CEO Guide to Disaster Readiness, and I ran through several of the pieces as a practitioner: EMT, Hurricane Katrina responder, Velocity co-chair, technologist on the ground. ### The state of the art vs. the state of the practice I had been driving through the aftermath of Hurricane Katrina, delivering temporary secure shelters for emergency supplies. American Red Cross workers were navigating with Google Maps, but the maps were stale. Roads shown as passable had been washed away. Directions led to dead ends. > "One of the interesting things with being a pretty senior technology person operating in a disaster is that you get to see the state of the art versus the state of the practice." The gap was not a technology gap. The tools existed. The problem was that sophisticated tooling designed for stable environments, by people who had never been in a disaster zone, fell apart the moment it hit real-world conditions. "Frequently, you'd be working with them and they'd give you directions over closed streets or places that didn't exist any longer," I told them about the Red Cross workers using Google's mapping tools. ### Collaborative infrastructure for crisis BusinessWeek placed me alongside a larger group of technologists building disaster-resilient tools. GeoCommons, OpenStreetMap, and Mapufacture were all working on collaborative mapping systems that could be updated by anyone with knowledge of a place, letting communities correct the record in real time as conditions changed. These were not abstract civic projects to me. They were operational infrastructure problems. Reliable maps during a flood were a version of the same challenge I had been solving at Amazon: how do you build systems that degrade gracefully under load, recover quickly from failures, and stay useful in the worst conditions. Emergency response just made the stakes undeniably human. ### Making Maps Work When Disaster Strikes - BusinessWeek · Article · 2008-07-07 - Original: http://www.businessweek.com/technology/content/jul2008/tc2008076_867685.htm - Canonical: https://jesserobbins.com/mentions/making-maps-work-disaster-strikes-businessweek/ > "One of the interesting things with being a pretty senior technology person operating in a disaster is that you get to see the state of the art versus the state of the practice." > — Jesse Robbins Rachael King's piece in BusinessWeek's CEO Guide to Disaster Readiness on the failure of mapping tools in the Katrina aftermath. Jesse Robbins had to get across U.S. Route 90 quickly. Hurricane Katrina had wrecked the Gulf Coast, and Robbins, an emergency medical technician, was on a mission, delivering temporary secure shelters for emergency supplies. American Red Cross workers guided him using Google mapping tools in areas where street signs had been washed away. He hit a dead end. The passage had been destroyed by the storm. His route was based on dated imagery, not a live satellite feed. "Frequently, you'd be working with them and they'd give you directions over closed streets or places that didn't exist any longer," Robbins said of the Red Cross workers relying on Google's maps. Emergency responders encounter this problem constantly when disasters strike. It also explains what drives the companies building mapping tools designed to help people navigate the aftermath of floods, earthquakes, and other catastrophes. ### Collaborative Mapping GeoCommons runs a site where users can explore a large atlas of maps with various data and add their own information. "The advent of user-contributed data allows nontechnical people to publish their own maps," says Sean Gorman, CEO and founder of FortiusOne, which runs GeoCommons. After devastating flooding in the Midwest in May 2008, people created their own maps of bridge closures, flood zone outlines, and Home Depot locations where people could get supplies. The maps were made available to anyone. OpenStreetMap is a freely available map that lets anyone with knowledge of a place contribute from anywhere. Mikel Maron saw the potential for collaborative mapping in disasters and brought the idea to the U.N. Joint Logistics Center group responsible for making maps for first responders. The U.N. started testing the approach. ### Art vs. Practice Other mapping tools place limits on who can make updates. "We may not want to rely on the crowd for data in an emergency, so there are tweaks to the model possible," Maron said during a presentation about disaster tech at a conference in Burlingame, California. "A smaller crowd of people who have some measure of responsibility in a situation should be able to pass along information." After Cyclone Nargis hit Myanmar, the UNJLC used maps that collected information about flooding and the state of critical transportation and health infrastructure, asking the community to email updates about roads, bridges, ports, and waterways. Robbins' work in Louisiana after Katrina showed him firsthand that some tools don't work nearly as well on the ground as they do in a lab. > "One of the interesting things with being a pretty senior technology person operating in a disaster is that you get to see the state of the art versus the state of the practice." He's hoping that with each disaster, he sees less of a difference between the two. ### Understanding Operations Culture (Part 1) - O'Reilly Radar · Article · 2008-06-14 - Original: http://radar.oreilly.com/2008/06/web-operations-culture-part1.html - Canonical: https://jesserobbins.com/mentions/understanding-web-operations-culture-part-1-oreilly-radar/ > "You don't choose the moment, the moment chooses you. You only choose how prepared you are when it does." > — Fire Chief Mike Burtch A 2008 O'Reilly Radar essay defining web operations culture through lessons drawn from the fire service. *The original URL is no longer live. This post is preserved from the O'Reilly Radar archive.* "You don't choose the moment, the moment chooses you. You only choose how prepared you are when it does." — Fire Chief Mike Burtch *(Note: I became a Firefighter-1 and EMT in 2000. My experiences in the fire service profoundly influence my efforts in technology. Much of my work over the past few years has been translating and distilling my knowledge from these two worlds, teaching others, and finding ways to apply it in the service of both.)* Last week I came upon a truck vs. scooter accident on my way home. I could hear a woman yelling in pain from underneath the truck (a good sign!) and could see a guy in the cab looking panicked and touching his controls. I stopped my car and "surveyed the scene" looking for things that might kill me (traffic, hazmat, downed power lines) or make the situation worse if undetected (additional victims, deflating tires, fires). It looked like the driver was about to move his truck, which would have definitely made things worse. I used my "command voice" to yell "Put it in park! Stop your engine! Set your brake! Get out and wait!" as I approached the truck. A city crew came over, and one of them told me "We've called 911 and they are on their way." I asked them to handle traffic control as I approached my patient. I then introduced myself and asked her if I could help. (I have to obtain consent before assisting an injured person, and a response means I know they still have their Airway, Breathing, and Circulation intact.) Her legs were entangled in her scooter which was trapped underneath the truck. While she probably had broken her leg, it didn't look all that bad. She was still wearing her helmet and it wasn't seriously damaged which meant her head was probably okay too. I did a quick check for bleeding and other serious injuries and did a "mental status check" by asking her name, where she was ("on my way to school"), and what had happened ("I was riding and that a**hole RAN OVER ME!"). This meant she was alert and oriented, which was good. Now that I was sure there weren't any other life-threatening injuries, I prepared to hold her head for c-spine stabilization. (Once you start holding stabilization, you cannot move again until you are ready to put the patient on a backboard.) As I positioned myself on the ground and took hold of her head, I explained: "I'm going to hold your head now to protect your neck and back. Once the fire department gets here, they are going to get your legs unstuck and then we'll get you on a backboard. Your job is to keep still and keep talking to us. There will be a lot of commotion and noise around you, and that's okay. Everyone will be watching out for you and so there is no reason to be scared. We've got you." As the fire department arrived they too surveyed the scene and I gave my quick report to the medics. They freed her legs and we transferred her to a backboard. I was released from the scene just as they started removing her helmet, and never even saw her face. Why am I telling you this story? I'm telling this story to illustrate how Operations culture works and to provide a little insight into how it is created. The city workers showed up, called 911, and made it safe for me to treat the patient by controlling traffic. I stopped the truck driver from further injuring the patient and stabilized her until the fire department and medics arrived. The medics took her to the hospital ER where she was probably treated and released. This is exactly how things should have gone in this situation. It happened because of people with a common desire or duty to act, training on how to act, and experience actually doing it. This is the essence of effective Operations culture. What does this have to do with Web Operations? Organizations that depend on the web will die if their site crashes and they don't recover. The longer the outage, the worse the damage often is. The same kind of Operations culture is required to effectively respond to, recover from, and prevent outages. While this seems obvious for many people with years of experience working on the web, it is a significant and often difficult shift for those in the mainstream. This seems particularly true for executives who think of Web Operations as an extension of corporate IT. This gap becomes especially painful when people accustomed to traditional "command-and-control" management styles and models try to apply them to this new type of organization. The CEO cannot shout or fire the website back up. The CFO cannot account, control, or audit the website back up. The Chief Counsel cannot sue it back to life. The CMO and their entire marketing team will not spin a website back online. The CIO or CTO probably can't recover the site either, at least not very quickly. The fate of the company frequently and acutely rests in the hands of engineers who do Web Operations. In thinking about how I got there — to a place where I could just act — I realized there was a lot to say. These were skills that took years of study, practice, and experience to develop. They weren't obvious. In some cases they were counter-intuitive. In many cases they required fighting your instincts. They required an enormous amount of trust in the people around you, combined with a large amount of independent judgment and action. They required clear communication under stress. Most importantly, they required care — about the outcome, about the people involved, and about doing the work well. I am proud of what the technology industry has accomplished. We have created incredible systems that have transformed society. Yet, I see a huge gap between what we know how to do and the actual state of practice in web operations. I meet people every day who are living by the seat of their pants, spending enormous amounts of time doing work that is well below their capability. I know exactly how hard it is to bring sanity to chaotic situations, and I want to share what I have learned, in hopes that it can help. ### Operations Is a Competitive Advantage (Secret Sauce for Startups!) - O'Reilly Radar · Article · 2007-10-23 - Original: https://www.oreilly.com/pub/a/radar/2007/10/operations-is-a-competitive-adv.html - Canonical: https://jesserobbins.com/mentions/operations-competitive-advantage-oreilly-radar/ My 2007 O'Reilly Radar argument that operations is a competitive advantage for startups, and occasionally a strategic weapon. Comments thread includes Luke Kanies, John Allspaw, John Willis, and Steve Loughran.

A note from Jesse

This post started it all. I wrote it for O'Reilly Radar in October 2007 after a summit where I kept making the same argument over lunch: operations is not a cost center. It is a competitive advantage and occasionally a strategic weapon. The original URL redirects now, so I am preserving it here. I had no idea what this post would set in motion. This was the start of Velocity. Luke Kanies introduced me to Adam Jacob in the comments and we started Chef.
My lunchtime conversations at the Summit centered around operations as a competitive advantage (and occasionally a "strategic weapon"). This advantage is the ability to consistently create and deploy reliable software to an unreliable platform that scales horizontally. Many people think of operations as "a bunch of boring work... which I'm hoping someone else is doing." It often takes less time to set up a development environment than the tools and infrastructure needed to test, deploy, monitor, and scale new software. The survival of most projects depends on working software, at least initially, and so if there is money or time many people will spend it on development. Unfortunately, people say they will "figure that ops stuff out soon", but what they mean is "when we're totally screwed!!!" It doesn't have to be that way. ![Side-by-side chart comparing Traditional Operations versus Secret Sauce Operations over 12 weeks, showing dramatic reduction in manual hours when infrastructure is automated early](/images/mentions/ops-roi-jesserobbins.png) Consider two Web 2.0 startups scaling to 20 systems during their first three months. The first team starts writing software and installing systems as they go, waiting to deal with the "ops stuff" until they have an "ops person". The second team dedicates someone to infrastructure for the first few weeks and ramps up from there. They won't need to hire an "ops person" for a long time and can focus on building great technology. In my experience it takes about 80 hours to bootstrap a startup. This generally means installing and configuring an automated infrastructure management system (Puppet), version control system (Subversion), continuous build and test (frequently CruiseControl.rb), software deployment (Capistrano), and monitoring. Once this is done the "install time" is reduced to nearly zero and requires no specialized knowledge. This is the first ingredient in "Operations Secret Sauce". This kind of scalability becomes really interesting when you find yourself suddenly popular, as iLike did when it launched its Facebook app and had to scale up fast. From their blog: > In our first 20 hours of opening doors we had 50,000 users sign up, and it is only accelerating. (10,000 users joined in the first 12 hrs. 10,000 more users in the next 3 hrs. 30,000 more users in the next 5 hrs!!) > > We started the system not knowing what to expect, with only 2 servers, but ready with backup. Facebook's rabid userbase chewed up our 2 servers almost instantly. We doubled our capacity to catch up. And then we doubled it again. And again. And again. Oh crap, we ran out of servers!! > > We just emailed everybody we know across over a dozen Bay Area startups, corporations, and venture firms in a desperate plea to find spare servers so we can triple our capacity for the continued onslaught. Tomorrow we are picking up over 100 servers from different companies to have them installed just to handle the weekend's traffic. Not being able to acquire hardware fast enough is by far a better problem than not being able to install it. *Are any VCs out there including effective operations in their due diligence? Are startups incorporating this in their pitch?* **Update:** Luke Kanies points out Adam Jacob's post about implementing Puppet for iLike. (Disclosure: I'm discussing collaboration with Adam's company, HJK Solutions.) Adam wrote: > Puppet enables us to get a huge jump-start on building automated, scaleable, easy to manage infrastructures for our clients. Using Puppet, we: > > 1. Automate as much of the routine systems administration tasks as possible. > 2. Get 10 minute unattended build times from bare metal. Puppet takes it the rest of the way. It's down to two and a half minutes for Xen. > 3. Bootstrap our clients' production environments while building their development environment. Because we are expressing the infrastructure at a higher level, when it comes time to deploy your production systems, it's really a non-event. > 4. Cross-pollinate between clients with similar architectures. When we solve a problem for one client, we've effectively solved it for the others. Sounds good to me. **Update #2:** John Allspaw of Yahoo/Flickr fame has great commentary on procurement and capacity management challenges for successful startups. --- ## From the original comment thread *Several of these people would go on to shape the DevOps movement.* **John Willis** wrote: There are some really interesting companies working around S3/EC2. Now you add a little Puppet with elastic computing cloud with a touch of services, now you got yourself a VC play. **Luke Kanies**, creator of Puppet, wrote: This article isn't clear on it, but iLike actually used Puppet to scale with all of those new machines. (I'm the author of Puppet, but iLike isn't a client of mine and I'm not the author of that post.) **John Allspaw**, Yahoo/Flickr, wrote: IMHO, bringing up ops considerations early in the process of product design should be a priority. Having at least some awareness of operational constraints can pay off significantly later on. I might even add that while not being able to acquire hardware fast enough is by far a better problem, streamlining your procurement process should also be considered part of the whole system. We've had to learn a bit of that at Flickr, just like iLike. :) As usual: great post, Jesse. **Steve Loughran**, SmartFrog team, wrote: Deployment should be part of every project, be it startup or in-house. If you can bring up a system and deploy to it during the build, then you get to do functional testing on it. If your app's diagnostics are in a form that the ops team can use and understand, then everyone benefits. I think where startups are special is there is less of a barrier between ops and dev; ideally: none. **Peter van Hardenberg**, founder of 3Tera, wrote: A software startup should be able to focus on the main value they provide and how to sell it better. Note to VCs: how long are you going to have the same 1MM or so spent in each startup for the same thing, namely operations? Find a way to not reinvent the wheel for each company. ### You Become What You Disrupt - O'Reilly Radar · Article · 2007-10-02 - Original: https://web.archive.org/web/20130228024606/http://radar.oreilly.com:80/2007/10/you-become-what-you-disrupt.html - Canonical: https://jesserobbins.com/mentions/you-become-what-you-disrupt-oreilly-radar/ > "You become what you disrupt. What changes occur when you win a platform play, when you go from disruptive technology to a public utility?" > — Jesse Robbins What happens when a disruptive technology wins a platform play and inherits the obligations of the system it replaced? I wrote this on O'Reilly Radar in 2007. The questions are still open.

A note from Jesse

I originally posted this to [O'Reilly Radar](https://web.archive.org/web/20130228024606/http://radar.oreilly.com:80/2007/10/you-become-what-you-disrupt.html) on October 2, 2007, in advance of the Web 2.0 Summit.
An idea we've been exploring in advance of the Web 2.0 Summit is "You become what you disrupt": 1. What changes occur when you win a platform play, when you go from disruptive technology to a public utility? 2. Where are the opportunities to innovate instead of regulate? 3. What parts of "eTel" are becoming "Tel"? 4. Where else will this happen? For example, the ongoing US VoIP/911 debacle was a missed opportunity to improve a life-saving technology. There have been years of delays, lawsuits, and regulatory standoffs while emergency calls go unanswered. Yes, the fine print says that VoIP isn't a replacement for your phone line, and suggests that you educate "anybody that might be in your home" about how to call 911. At some point in the adoption curve that kind of disclaimer becomes unacceptable... *but where?* The 911 Modernization and Public Safety Act (H.R. 3403) was being considered by Congress. This bill was intended to give VoIP providers the same access to the 911 system as wireless carriers. It had broad support from both public safety officials and VoIP providers, but was opposed by established operators because it "provides more access to 911 infrastructure than wireless carriers have and therefore an unfair advantage". How do we avoid this kind of problem in the future? Where should we be looking now? Services like Skype are classified as "data services", meaning they don't have to provide 911 access for now. It's unclear how services like SkypeOut and Skype embedded handsets change this, although their terms of service say: > "*7.4.2 No Compulsion to Offer Emergency Services.* You recognize and agree that Skype is not required to offer Emergency Services pursuant to any applicable local and or national rules, regulation or law. You further recognize that Skype is not a replacement for Your primary telephone service." Perhaps this Skype job posting provides clues as to how this will play out: > **Director of Government and Regulatory Affairs, North America** > > - Influence legislative and regulatory developments in the North American region > - Minimize exposure to political and regulatory risk > - Develop specific expertise in the area of public safety and state telecommunications public policy > - Promote Skype's interests through various coalitions, trade associations, and public safety groups > - Act as early warning system for regulatory risks These questions take us far beyond 911. The Telecommunications Act of 1996 states that "consumers in all regions of the nation, including low income consumers should have access to telecommunications and information services". Will the Universal Service Fund subsidize internet and VoIP? Should it? Similarly, new "utilities" are emerging from the web as a platform. Will utility computing services like Amazon EC2 eventually become regulated? What about identity services, or even "social utilities" like Facebook? ## Patents ### Integrating logic services with group communication services (US11908471B2) - Granted 2024-02-20 · Orion Labs, Inc. - https://patents.google.com/patent/US11908471B2 System and methods for integrating logic services with group communication services, including voice assistant capabilities. Enables programmable workflows and intelligent automation within real-time group voice communication channels. ### Bot group messaging method and system (US11711326B2) - Granted 2023-07-25 · Orion Labs, Inc. - https://patents.google.com/patent/US11711326B2 Methods and systems for bot group messaging in a group communication platform. The system enables automated bot participants to interact within group voice and messaging channels, processing natural language inputs and generating contextual responses for team communication workflows. ### Transcription bot for group communications (US11258733B2) - Granted 2022-02-22 · Orion Labs, Inc. - https://patents.google.com/patent/US11258733B2 A group communication service receives user node communications from and distributes user node communications to members of a communication group. The service receives an audio transcription request from a transcription bot node and provides transcription of group communications. ### Group communication device providing bypass connectivity (US11128521B2) - Granted 2021-09-21 · Orion Labs, Inc. - https://patents.google.com/patent/US11128521B2 A group communication device and method providing bypass connectivity that enables communication even when primary network infrastructure is unavailable. The device establishes alternative communication paths for resilient group operations. ### Operating environment partitioning for securing group communication device resources (US10887290B2) - Granted 2021-01-05 · Orion Labs, Inc. - https://patents.google.com/patent/US10887290B2 Methods and systems for partitioning operating environments to secure resources on group communication devices. The approach isolates sensitive communication data and system resources to prevent unauthorized access across application boundaries. ### Transcription bot for group communications (US10855626B2) - Granted 2020-12-01 · Orion Labs, Inc. - https://patents.google.com/patent/US10855626B2 A transcription bot system for group communications that automatically converts spoken voice messages into text transcriptions within group communication channels, enabling searchable records and accessibility for team conversations. ### Intelligent agent features for wearable personal communication nodes (US10462003B2) - Granted 2019-10-29 · Orion Labs, Inc. - https://patents.google.com/patent/US10462003B2 Systems, methods, apparatus and software enable intelligent agent features for user nodes that are members of a communication group. Instructions instantiate an intelligent agent node as a member of the communication group. ### Dynamic muting audio transducer control for wearable personal communication nodes (US10120644B2) - Granted 2018-11-06 · Orion Labs, Inc. - https://patents.google.com/patent/US10120644B2 Dynamic muting and audio transducer control methods for wearable personal communication nodes. The system intelligently manages audio output and muting behavior based on environmental context and user activity to optimize the communication experience on wearable devices. ### Intelligent agent features for wearable personal communication nodes (US10110430B2) - Granted 2018-10-23 · Orion Labs, Inc. - https://patents.google.com/patent/US10110430B2 Intelligent agent features integrated into wearable personal communication nodes. The system provides voice-activated intelligent agent capabilities on wearable devices, enabling hands-free access to information, task completion, and contextual assistance within group communication workflows. ### One-touch group communication device control (US10057394B2) - Granted 2018-08-21 · Orion Labs, Inc. - https://patents.google.com/patent/US10057394B2 A one-touch control system for group communication devices that simplifies device operation to a single physical interaction, enabling instant access to group voice channels for frontline and field workers. ### Wearable communication device (USD816051S1) - Granted 2018-04-24 · Orion Labs, Inc. - https://patents.google.com/patent/USD816051S1 The ornamental design for a wearable communication device, substantially as shown and described. ### Device to device grouping of personal communication nodes (US9936010B1) - Granted 2018-04-03 · Orion Labs, Inc. - https://patents.google.com/patent/US9936010B2 Methods and systems for device-to-device grouping of personal communication nodes. Enables dynamic formation and management of communication groups across wearable devices without requiring centralized infrastructure. ## Topics - DevOps (26 items): https://jesserobbins.com/topics/devops/ - Venture Capital (18 items): https://jesserobbins.com/topics/venture-capital/ - Developer Tools (15 items): https://jesserobbins.com/topics/developer-tools/ - Chef (12 items): https://jesserobbins.com/topics/chef/ - Engineering Culture (12 items): https://jesserobbins.com/topics/engineering-culture/ - Site Reliability Engineering (12 items): https://jesserobbins.com/topics/site-reliability-engineering/ - AI Developer Tools (11 items): https://jesserobbins.com/topics/ai-developer-tools/ - Resilience Engineering (11 items): https://jesserobbins.com/topics/resilience-engineering/ - Amazon (10 items): https://jesserobbins.com/topics/amazon/ - Seed Stage Investing (9 items): https://jesserobbins.com/topics/seed-stage-investing/ - Chaos Engineering (9 items): https://jesserobbins.com/topics/chaos-engineering/ - AI Infrastructure (8 items): https://jesserobbins.com/topics/ai-infrastructure/ - Web Operations (8 items): https://jesserobbins.com/topics/web-operations/ - Open Source (7 items): https://jesserobbins.com/topics/open-source/ - Incident Response (7 items): https://jesserobbins.com/topics/incident-response/ - Infrastructure as Code (6 items): https://jesserobbins.com/topics/infrastructure-as-code/ - Velocity Conference (6 items): https://jesserobbins.com/topics/velocity-conference/ - Awards (5 items): https://jesserobbins.com/topics/awards/ - Heavybit (5 items): https://jesserobbins.com/topics/heavybit/ - Cloud Infrastructure (5 items): https://jesserobbins.com/topics/cloud-infrastructure/ - DevOps History (5 items): https://jesserobbins.com/topics/devops-history/ - Early-Stage Investing (4 items): https://jesserobbins.com/topics/early-stage-investing/ - AI (4 items): https://jesserobbins.com/topics/ai/ - Business Insider (3 items): https://jesserobbins.com/topics/business-insider/ - Investor Rankings (3 items): https://jesserobbins.com/topics/investor-rankings/ - Infrastructure (3 items): https://jesserobbins.com/topics/infrastructure/ - AI Investments (3 items): https://jesserobbins.com/topics/ai-investments/ - Incident Management (3 items): https://jesserobbins.com/topics/incident-management/ - Operations (3 items): https://jesserobbins.com/topics/operations/ - Wired (3 items): https://jesserobbins.com/topics/wired/ - GameDay (3 items): https://jesserobbins.com/topics/gameday/ - Seed Stage (2 items): https://jesserobbins.com/topics/seed-stage/ - Founders (2 items): https://jesserobbins.com/topics/founders/ - Agentic Developer Experience (2 items): https://jesserobbins.com/topics/agentic-developer-experience/ - Software Engineering Evolution (2 items): https://jesserobbins.com/topics/software-engineering-evolution/ - Developer Platforms (2 items): https://jesserobbins.com/topics/developer-platforms/ - Data Pipelines (2 items): https://jesserobbins.com/topics/data-pipelines/ - Enterprise AI (2 items): https://jesserobbins.com/topics/enterprise-ai/ - Startups (2 items): https://jesserobbins.com/topics/startups/ - GameDay Testing (2 items): https://jesserobbins.com/topics/gameday-testing/ - O'Reilly (2 items): https://jesserobbins.com/topics/o-reilly/ - Orion Labs (2 items): https://jesserobbins.com/topics/orion-labs/ - OnBeep (2 items): https://jesserobbins.com/topics/onbeep/ - Voice Communication (2 items): https://jesserobbins.com/topics/voice-communication/ - Wearable Technology (2 items): https://jesserobbins.com/topics/wearable-technology/ - Master of Disaster (2 items): https://jesserobbins.com/topics/master-of-disaster/ - Firefighting (2 items): https://jesserobbins.com/topics/firefighting/ - Infrastructure Automation (2 items): https://jesserobbins.com/topics/infrastructure-automation/ - Cloud Computing (2 items): https://jesserobbins.com/topics/cloud-computing/ - Disruption (2 items): https://jesserobbins.com/topics/disruption/ - Disaster Response (2 items): https://jesserobbins.com/topics/disaster-response/ - Humanitarian Tech (2 items): https://jesserobbins.com/topics/humanitarian-tech/ - Emergency Management (2 items): https://jesserobbins.com/topics/emergency-management/ - OpenStreetMap (2 items): https://jesserobbins.com/topics/openstreetmap/ - Civic Tech (2 items): https://jesserobbins.com/topics/civic-tech/