The gap between building fast and shipping safely

· 6 min read
AI-generated image: The gap between building fast and shipping safely
AI-generated image

The productivity-software news on 4 August circled a single tension: the tools got faster at producing work, and the hard part moved to trusting that work in daily use. A guided-selling playbook that builds in two minutes, code an agent writes in one pass, an agent that needs steady access to your apps — each headline was really about the distance between something working once and something working every day.

That distance is the useful lens for a business choosing tools right now. A demo shows you the ceiling. What you are buying is the floor: what the tool does on an ordinary Tuesday, with your data, in the hands of people who did not build it. The announcements below are ordered by how much they change that decision.

The distance between a demo and a deployment

The clearest statement of the day came from a write-up of a live demo at a SaaStr event, where the quote-to-cash vendor Nue showed a guided-selling playbook being assembled on stage. The headline carried the whole lesson: the playbook took two minutes to build, and the full implementation still takes ninety days [1]. The account notes that the presenter let the agent hit its limits on screen, on conference Wi-Fi, in front of an audience — which is the honest way to show a tool.

The two numbers are not in tension; they describe two different jobs. Building is the cheap part now. What fills the ninety days is everything that makes the built thing correct: connecting it to real records, encoding the rules your business actually follows, handling the cases the demo skipped, and getting people to change how they work. When you evaluate any AI-assisted tool, ask the vendor to separate those two timelines out loud. A two-minute build with an unstated implementation cost is not a saving; it is a deferred bill. We have written before about the real cost of "we'll build it later", and the same arithmetic applies when the thing being deferred is configuration rather than code.

Giving AI agents dependable access to your apps

A piece from Zapier put its finger on why so many agent projects stall: "As advanced as AI agents have become, they still fall short in one respect: connecting to your apps reliably" [2]. The reason it gives is mundane and correct — every app authenticates, structures its endpoints and shapes its data in its own way, and each breaks in its own way too.

This is the part of the AI story that does not make headlines and decides whether the project survives. An agent that can reason well but cannot reliably read a record, write a note, or trigger the next step is a clever assistant with its hands tied. The friction is not intelligence; it is plumbing. Before you commit to an agent-based workflow, map which systems it must touch and how solid each connection is. A connection that works in a demo but times out under load, or silently drops a field, will produce confident, wrong actions. This is the same problem we described in why your tools do not talk to each other: the value of any tool is capped by how well it reaches the systems around it, and that ceiling is set long before the agent writes a word.

Making AI-written code reviewable

GitHub published guidance aimed at teams whose coding agents produce large changes at once. Its advice: "Instead of one huge, un-reviewable pull request, teach coding agents to decompose work into a clean, ordered stack" of smaller, ordered changes [3].

The principle travels well beyond code. When a machine can generate a large amount of output quickly, the bottleneck moves to review. A single enormous block of AI-produced work — a contract draft, a migration plan, a bulk data change — is hard to check, so people check it less, and errors ride through. Breaking the work into an ordered sequence of smaller pieces, each of which a person can actually read and approve, is how you keep human judgement in the loop without slowing to a crawl. If your team is starting to let AI produce work at scale, decide in advance how that work gets broken up for review. Our note on how to tell whether a task should be automated covers the companion question: some outputs need a person on every one, and some do not, and knowing which is half the discipline.

Building tools without writing code

GitHub also described how its own legal team used a command-line AI tool to streamline day-to-day work, under a plain promise: "Learn how to build tools to simplify how you work—without writing a single line of code" [4].

The detail worth noticing is who did it. Not the engineering team — the legal team. When the people who understand a process most closely can assemble their own small tools, the work no longer waits in an IT queue and does not lose meaning as it is handed between people. That is a real shift in who gets to automate. It also raises a governance question every business should answer before it spreads: who owns and maintains the tools that non-technical teams build, and what happens when the person who made one moves on. The upside is speed and accuracy at the source; the trade-off is a wider surface of tools that someone still has to stand behind.

Choosing a starter stack

For earlier-stage companies still assembling their toolset, Intercom published a guide to a foundational stack, framed as helping you "create a foundational tech stack for your high growth and early-stage startup" [5]. Any such list is a starting point rather than a prescription, and it is most useful read that way.

The value of a curated list is that it narrows the field, which is worth a lot when every category has dozens of credible options. The risk is treating someone else's list as your decision. A stack that suits one company's motion can be wrong for yours, and switching costs compound as you grow into a tool. Use a list like this to shorten your shortlist, then judge each entry against your own work: what it must connect to, who will run it, and what it costs to leave. That is the method we set out in choose software worth using — the tool that fits your habits beats the tool that tops a ranking.

Securing the agents themselves

Underneath all of the above sits a question the industry is only starting to organise around. A week after its formation, the Open Secure AI Alliance — described as "spearheaded by Nvidia and grown to over 120 companies" — already has proposals out for defending against AI agents [6].

That this is happening at industry scale tells you something. As agents gain the ability to act — to reach into your apps, move data and trigger steps — they become a surface to defend, not just a feature to enjoy. A business adopting agents this year does not need to solve this alone, but it does need to ask its vendors what an agent can and cannot do inside its systems, and how that boundary is enforced. The convenience and the exposure grow from the same root: an agent is useful exactly to the degree that it can act, and worth watching for the same reason.

The thread, pulled together

Every item today described the same second half of the AI story. The building is quick; the trust is slow. Reliable connections, reviewable output, clear ownership of home-made tools, a stack chosen for your own work, and a boundary around what agents may do — none of these show up in a demo, and all of them decide whether the demo becomes something you run. When you are choosing tools this month, spend your scrutiny there. The floor matters more than the ceiling.

Sources

  1. [1] Nue's Guided Selling Playbook Took 2 Minutes to Build. The Full Implementation Still Takes 90 Days. — SaaStr
  2. [2] How to give your AI agents reliable app access for free — Zapier
  3. [3] Turn one giant AI-generated pull request to a reviewable stack — The GitHub Blog
  4. [4] How the GitHub legal team used Copilot CLI to streamline their workflows — The GitHub Blog
  5. [5] The 9 best tools for your early-stage startup tech stack in 2026 — Intercom
  6. [6] Nvidia doesn't mess around: A week after open AI industry group formed, it's already showing progress — TechCrunch

The 360REV newsletter

What is actually changing across productivity software, written for operators and cited to sources. No more than one email a day.

Double opt-in — we send one confirmation link and nothing else until you click it. Unsubscribe from any edition. We never sell or share your address.