Skip to content
ANALYSIS

The difference between experimentation and operations

Why a successful sandbox is only the beginning, and which governance, engineering and operating evidence production requires.

By Cytria Research3 min read

Standfirst. A sandbox proves that a task may be possible under selected conditions. Operations must make the workflow repeatable when data, users, vendors, errors and obligations change. The gap is governance and engineering work, not merely model tuning.

What a prototype proves

A prototype can test whether a model produces a useful output from a curated input. It can expose failure modes and help users refine requirements. It does not establish production reliability, lawful use, secure integration, supportability or an acceptable total cost.

The calendar’s “20%” should be treated as a rhetorical illustration, not a measured universal ratio. The production share varies by workflow and existing capability.

What operations adds

  1. Defined service: purpose, users, supported inputs, exclusions and service owner.
  2. Information controls: classification, provenance, permissions, retention and deletion.
  3. Architecture: identity, integration, network, keys, logging, resilience and exit.
  4. Assurance: representative tests, acceptance thresholds, red-team scenarios and independent review proportionate to risk.
  5. Human control: checkpoints, evidence, exceptions and separation of draft from action.
  6. Operations: monitoring, incidents, support, vendor and model changes, rollback and retirement.
  7. Economics: licence, integration, review, exception, assurance and change costs.

The production gate

Question Evidence required
Is the outcome useful? Baseline and acceptance test
Is use authorised? Recorded legal, privacy and business review
Can access be controlled? Role and negative-permission tests
Are failures bounded? Scenarios, thresholds and stop mechanism
Can it be supported? Owner, runbook, monitoring and incident path
Can it change safely? Versioning, regression test and rollback
Can it end safely? Export, deletion and retirement plan

NIST frames AI risk management as continuous across Govern, Map, Measure and Manage. FINMA’s guidance similarly emphasises lifecycle governance, inventories, testing and monitoring in supervised finance. A successful demo addresses only a subset of that work.

Common transition errors

Teams carry prototype credentials into production, rely on a developer’s personal account, test only ideal prompts, omit review time from costs, or connect an agent to execution before exceptions are understood. Another error is indefinite “pilot” status: real users depend on the system while ownership and incident duties remain temporary.

Cytria’s operational interpretation

Experimentation answers “could this help?” Operations answers “who is responsible, under which conditions, with what evidence, and what happens when it fails or changes?” The transition is complete only when the second question has a maintained answer.

Limitations, sources and metadata

General operational information. The effort ratio must be estimated for the specific workflow rather than quoted as fact.

  • NIST, AI RMF Core and Generative AI Profile; FINMA, AI governance guidance; reviewed 14 July 2026.
  • Type: Analysis
  • Title tag: The difference between AI experimentation and operations | Cytria
  • Meta description: Why a successful sandbox is only the beginning, and which governance, engineering and operating evidence production requires.
  • Slug: `the-difference-between-experimentation-and-operations`
  • Author / owner: Cytria Research / Cytria
  • CTA: Assess one prototype against the production gate
  • Editorial risk: Do not present 20% as a measured or universal figure.

Recommended next step

Evaluate the most practical path to deploy controlled AI inside your business operations.

Start free diagnostic