On September 16, 2026, OpenAI published guidance on connecting ChatGPT Work and Codex activity to business value. The useful lesson is not a headline ROI percentage. It is a method: activity inside an AI tool should be compared with the outcome of the workflow it supports. The same principle applies to a custom application with AI integration.
Frequent use is not the same as a return
A team can send thousands of messages or spend more credits on a task without producing work faster, more accurately or more profitably. Conversely, a workflow used by only a few people may avoid costly errors. Interaction volume measures adoption, not value.
OpenAI describes Admin Console views for active users, credits and tokens, followed by a classifier that groups a sample of messages into use cases and tasks. Engineering views include measures of Codex contributions to merged code and review activity. Official documentation cautions against treating message share, credit share and token share as the same metric.
Those views indicate where a business question is worth asking. They cannot, by themselves, calculate the profit created by AI. Access can depend on workspace permissions and enabled features, and classification history starts only when classification begins.
Choose a specific workflow first
"We want AI ROI" is not a testable objective. A better question is whether the company can prepare quotations faster, reduce incorrectly resolved support tickets or process documents more efficiently without sacrificing accuracy. Choose one repeated task, one owner and an observation period.
Then establish a baseline. How many cases arrive each week? How long does a case take? How many require correction? What does an error cost? Without a before picture, a later improvement may be mistakenly attributed to AI when it came from seasonality, a new team member or another process change.

A calculation that resists cosmetic numbers
A starting model is: net benefit = value of genuinely released capacity + demonstrable additional margin - implementation cost - usage cost - review cost - losses caused by errors. Not every saved hour becomes revenue. If the released time is not used productively, report capacity created rather than cash earned.
Here is a hypothetical example. A company processes 200 requests each month, each previously taking 20 minutes. With AI, a first draft takes 7 minutes, but a human needs another 5 to verify it. The true saving is 8 minutes per request, about 26.7 hours per month, not 43.3. Subtract the application, integration and rework costs. The example illustrates a method, not a promised outcome.
Quality must be a threshold, not an afterthought. If the work becomes faster but produces more complaints or incorrect CRM data, the apparent saving has merely moved to another team. Compare the time from request to validated result, not just the time to the first AI answer.
Metrics should follow the process
- Support: resolution time, reopened tickets, satisfaction and handoff to a person.
- Sales: quotation preparation, qualified opportunities, win rate and margin, not just email volume.
- Operations: processing time, exceptions, rework and review time.
- Software engineering: delivery, defects, review time, rework and production quality.
ChatGPT Work or Codex usage data can start the investigation; business-system data must complete it. Connect the view to CRM, helpdesk, inventory, billing or repository outcomes. Without that link, the dashboard maps tool use rather than impact.
When an AI-integrated application is justified
A general-purpose tool may be enough when the task is occasional, the data is simple and a person can verify the result easily. A custom application with AI integration becomes useful when the workflow must connect to internal systems, enforce approvals or permissions, or handle repeated volume. A business dashboard with automated reporting can then bring the measurement together with operational results.
Expand only after a pilot with a defined target, time period and owner. If the result holds, standardize sources and checks. If not, change the workflow or stop. Sound measurement also means recognizing when AI is not worth adding.
What a four-week pilot can look like
In week one, select one activity and record a baseline: volume, time to a validated result, errors and cost. In week two, configure access to necessary information and define where a person decides. In week three, run the assisted process alongside normal review without changing other major rules. In week four, compare similar cases and speak with the people who used the tool.
Do not collapse the result into one average. If 80% of requests are straightforward and 20% are expensive exceptions, the average might hide that AI helps only with easy cases. Segment by request type, not just by user. Record how many tasks were handed to a person and how many answers required a complete rewrite.
The problem with an unexplained ROI percentage
OpenAI includes a hypothetical return calculation based on released employee capacity. Such a model can be useful, but assumptions about how many hours become productive work and what that work is worth must be stated. "We saved 100 hours" does not automatically mean the company earned the value of 100 hours or eliminated a cost. For services, check whether more projects were delivered without reducing quality; for sales, whether margins or won opportunities improved.
For the same reason, a case-study number should never be copied into a proposal for another business. Companies have different volumes, workflows, data and levels of control. A good result elsewhere is a hypothesis to test, not a contractual promise.
Governance is part of both cost and value
For an application connected to customer data, include permissions, action logs, the ability to correct or stop automation, and source maintenance in the evaluation. These are real costs, but also safeguards against more expensive mistakes. A project that ignores security to achieve a quick paper ROI may simply push risk into the future.