From prototype to accountable operation
OpenAI introduced Presence on 22 July 2026 as a managed enterprise product for deploying AI agents in voice and chat workflows. The announcement is significant less because it presents another general-purpose chatbot than because it addresses the difficult operational work required when an agent can interact with customer records, follow company policies and make approved changes in business systems.
Presence is intended for defined jobs rather than unconstrained automation. OpenAI describes examples ranging from customer-support and sales conversations to claims handling, human-resources queries and IT service requests. In these settings, the agent may retrieve information, update systems or complete pre-authorised actions, but its access and authority are meant to be limited to the requirements of the task.
That framing recognises an important distinction in enterprise AI. A model may demonstrate useful reasoning in a controlled trial, yet remain unsuitable for production if it cannot reliably apply an organisation’s current policies, use tools appropriately, protect sensitive information and hand a case to a person at the right moment. The practical unit of deployment is therefore not merely the model; it is the model plus the surrounding rules, integrations, testing and oversight.
A product built around controls
Presence brings together policies, standard operating procedures, scoped permissions, guardrails, approved actions, simulations, evaluations and escalation paths. Companies decide which rules should remain consistent across channels and which should vary for a particular workflow. This matters for organisations seeking a consistent customer experience while acknowledging that a voice call, a web chat and an internal support request can carry different risks and operating requirements.
The platform’s focus on permissions is particularly consequential. An agent that only answers questions presents a different risk profile from one that can alter an account, reset access, apply a billing adjustment or route a claim. By tying actions to pre-approved boundaries and human hand-offs, Presence is designed to turn the question of trust into a set of operational decisions: what data can be used, which tools can be called, which outcomes are allowed and which circumstances require review.
OpenAI also emphasises pre-launch testing. Simulated interactions and evaluators are intended to test whether an agent reaches the intended outcome, complies with policy, uses tools correctly and escalates cases appropriately. Such testing cannot guarantee flawless behaviour in the real world, but it can make a deployment’s intended standards explicit before it reaches customers or employees.
This structure is broadly aligned with the risk-management approach advocated by the US National Institute of Standards and Technology. Its AI Risk Management Framework describes risk management as a continuing activity involving governance, contextual mapping, measurement and management, rather than a one-off compliance check. For agents that take actions in live systems, ongoing assessment is as material as initial performance testing.
Continuous improvement is the central proposition
Presence’s most distinctive promise is its post-launch improvement process. OpenAI says production sessions, escalations and other quality signals can identify where an agent is failing or becoming outdated as products, policies and customer behaviour change. Codex is then used to investigate those signals and suggest updates, which teams can test against the production version and approve for a controlled rollout.
The approach reflects a basic operational reality. Policies are revised, product catalogues change, customer issues evolve and connected software is updated. A static agent may therefore become unreliable even if it performed acceptably at launch. A governed feedback loop can help organisations identify drift, test changes and preserve a record of who authorised them.
However, continuous improvement also creates its own governance challenge. An automated proposal is not the same as a validated policy update. Organisations will need to determine who owns evaluation criteria, who can approve a modified workflow, how changes are documented and how they can be reversed if results deteriorate. The quality of the underlying knowledge, integrations and operating procedures will remain decisive.
OpenAI cites its own English-language telephone support as an early deployment, saying the system resolves a substantial share of inbound issues without human assistance. It also identifies BBVA, SoftBank and IAG as organisations exploring or developing uses of the product. These examples should be read as early evidence of potential applications rather than proof that a common operating model will work across sectors. Banking, insurance, employee support and sales each have different requirements for auditability, privacy, error tolerance and human accountability.
Managed delivery limits the initial market
Presence is available only to eligible enterprise customers through a limited general-availability programme. It is not a self-service product. Deployments are led by OpenAI Forward Deployed Engineers, with support from selected global systems integrators.
This delivery model is a realistic acknowledgement that many high-value workflows cannot be connected safely through a simple configuration screen. The work may involve mapping processes, integrating business systems, designing escalation routes, preparing test sets and resolving exceptions. Managed implementation can reduce the gap between a promising demonstration and an agent that fits an organisation’s actual controls.
The trade-off is that Presence will initially be less accessible than a self-serve software platform. Availability depends on workflow suitability, customer readiness and deployment capacity, while commercial terms and technical details are set for individual engagements. For OpenAI, this makes Presence a services-intensive enterprise offering as well as a software product.
What enterprises should measure
For prospective users, the relevant question is not whether an agent sounds capable in a conversation. It is whether it improves a defined process while staying inside acceptable bounds. Measures should include task completion, policy adherence, correct tool use, escalation quality, customer or employee outcomes, error severity and the time required to identify and correct failures.
Enterprises should also test difficult cases deliberately: incomplete information, contradictory records, attempts to exceed permissions, sensitive requests and changes in policy. Human escalation should be treated as a designed feature, not evidence that the automation has failed. In many workflows, the most valuable agent may be one that resolves routine cases efficiently and recognises the limits of its authority.
Presence therefore represents a maturing view of agentic AI. Its success will depend less on a single model benchmark than on whether OpenAI and its customers can establish dependable operating systems around agents: constrained access, measurable behaviour, accountable approvals and improvement without uncontrolled change. For businesses moving beyond experiments, that is likely to be the real test of production readiness.
Sources
- Introducing OpenAI Presence — OpenAI
- OpenAI Presence — OpenAI Help Center
- OpenAI Presence — OpenAI
- AI RMF Core — National Institute of Standards and Technology
- OpenAI Introduces Product to Deploy AI Agents in Workflows — Bloomberg Law



