AI for CIOs · Independent decision intelligenceSource-backed reporting · No paid editorial rankings
CIO AI Review

An architecture-and-operations review for technology executives deciding how AI should enter the enterprise stack, which controls must follow it, and where vendor demonstrations leave material questions unanswered.

CIO briefings

An Agents API sandbox needs a workload-specific custody map

OpenAI says its public-beta Agents API runs an OpenAI-managed harness while allowing the customer to select an OpenAI-hosted sandbox, its own infrastructure, or an integrated partner environment. Those are different execution and custody arrangements, not interchangeable security labels. A CIO should decide where one workload's files, secrets, tools, artifacts, network paths, logs, and recovery obligations reside before treating a successful agent session as production approval.

Answer capsule

OpenAI says its public-beta Agents API runs an OpenAI-managed harness while allowing the customer to select an OpenAI-hosted sandbox, its own infrastructure, or an integrated partner environment. Those are different execution and custody arrangements, not interchangeable security labels. A CIO should decide where one workload's files, secrets, tools, artifacts, network paths, logs, and recovery obligations reside before treating a successful agent session as production approval.

What the source establishes

  • OpenAI announced the Agents API in public beta on September 10, 2026 and says its service hosts and maintains the agent harness.
  • The official announcement describes three compute choices: an OpenAI-managed sandbox, customer infrastructure, or a sandbox partner; it lists partner options with different VPC, file, secret, compute, cold-start, and cost profiles.
  • OpenAI says hosted sandboxes can run code, work with files, and produce artifacts, and may be configured with files, packages, skills, and plugins.
  • The API describes long-lived sessions, automatic context compaction, tool search, programmatic tool calling, and subagents; these are capabilities, not proof of a customer's exact permissions or control effectiveness.

Classify the workload before picking a sandbox

Name a single business service, its data classes, permitted users, service owner, recovery objective, record-retention duty, and allowed agent actions. An isolated source-code analysis with synthetic files poses a different custody problem from a production incident response that touches logs, customer records, privileged credentials, and deployment tools. OpenAI's product announcement describes a hosted harness and multiple compute locations, but the CIO must map the actual request, context, instructions, model access, input files, tool calls, network traffic, generated artifacts, and logs through every provider and subcontractor. A VPC deployment label alone does not prove where all processing and support access occur.

Record the chosen model and API version, hosted/own/partner execution choice, regional and contractual boundaries, data return path, use of vaults or secrets, storage lifecycle, and the operational team with the right to suspend a session. Require evidence for the specific option proposed, not a generic diagram assembled from features available in different environments.

Treat a session as a governed process, not a single request

The announcement emphasizes context management across long sessions, discovery and use of tools, and subagent coordination. A service-owner review needs a bounded task, maximum duration and spend, delegated identity, tool allow-list, source-data permissions, human approval gates, and an expiry condition. Test a tool that returns a stale or malicious page, a delegated agent asking for a broader secret, a subagent continuing after the parent task ends, and an interrupted run resumed against changed permissions. Verify whether the audit trail can reconstruct which agent, tool, file, and person caused every consequential operation. None of those behaviors is established for a buyer by the beta announcement.

Context compaction may preserve task state while changing what detail is visible to the model in later steps. Decide which instructions and facts must be persisted outside the conversational context as authoritative records, and when an agent should stop and ask for an updated source. Retained outputs should distinguish proposed change, human approval, actual action, test result, and rollback; a plausible final answer does not replace a change ticket or system log.

Compare architecture options on recoverability and exit

Review hosted, own-infrastructure, and partner execution against the same work sample. Measure not just latency and token cost but file transfer, storage, human supervision, egress, restart behavior, exception handling, incident response, and operational burden. Document who patches the environment, rotates keys, grants a plugin permission, investigates a data spill, preserves logs, and restores a partially completed workflow. An integration that is easy to start may still be expensive to support or migrate, while a self-managed environment shifts responsibilities to the internal platform team.

The release decision should say which workloads are permitted, which remain in a sandbox with synthetic data, and which cannot use this architecture without further evidence. Test termination and export of files, session state, evaluation cases, and audit records when a provider or model is changed. The CIO owns the architecture and service acceptance; information security, privacy, legal, procurement, and the business owner retain their own decisions.

Turn this source into a reviewable decision

For AI for CIOs, use this briefing as a dated decision record rather than a substitute for the source. Preserve OpenAI, Introducing the Agents API, the exact URL, the September 22, 2026 review date, the supported facts above, the editorial interpretation, the limitations, and any buyer-specific evidence. Link that record to the decisions most directly affected: Enterprise AI platform architecture; Identity and agent access; Operations and incident intelligence; AI portfolio economics. State whether the source changes the scope, evidence requirement, control, sequence, or only the language used to describe the decision.

Before action, name the accountable owner, affected population and workflow, exact offering or configuration, source data and rights, human decision point, exception and appeal path, complete cost, expected benefit, failure and stop conditions, retained evidence, and next review date. Keep official facts, provider statements, buyer observations, representative tests, measured outcomes, editorial inferences, and unknowns visibly separate. Reopen the record when the source, offer, model, integration, data, policy, population, responsible person, or measured result changes.

Limitations and unknowns

OpenAI's September 10, 2026 announcement, reviewed September 22, describes a public-beta product and provider-stated architecture options. It does not prove a buyer's configured environment, exact data residency, encryption and deletion behavior, partner terms, sandbox isolation, tool authorization, secrets handling, audit completeness, session recovery, compatibility, uptime, performance, cost, or production suitability. The publication did not run a test deployment. Obtain current technical, security, legal, and contractual documents and exercise a representative controlled workload with the actual identity and support owners. This source predates the September 21, 2026 cutoff; no new post-cutoff launch is claimed.

Decision test

Ask whether the source changes the decision itself, the evidence required, the implementation sequence, or only the language used to describe an existing capability. Record which claims are directly supported, which are provider statements, which require an independent test, and which remain unknown. A source-linked review should make uncertainty easier to see, not bury it inside a blended score.

Questions to take into review

  • Which services are common and which remain workload-specific?
  • How can a team change a model without rewriting the application?
  • Whose authority is the agent exercising?
  • Can each tool call be attributed and reversed?
  • Which telemetry is missing or sampled?
  • Can the model change production or only advise?
  • What is the unit of useful work?
  • How does cost change with context, retrieval, tool calls, retries, and review?
The publication supports research and executive decision preparation. It does not provide legal, financial, accounting, employment, clinical, cybersecurity, investment, procurement, or implementation advice.