OpenAI / 6 October 2026

Ironclad Research Shows Gains and Limits for OpenAI Agents

OpenAI disclosed research with contracting software company Ironclad on 6 October in which GPT-6 Astra scored 55.0% against task criteria, compared with 41.6% for GPT-5.6 Sol. The evaluation covered 11 legal, commercial and procurement workflows. Estimated time per attempt fell from 37 minutes to 19.2 minutes, but those timings were simulated estimates rather than measured customer savings. The collaboration matters because it tests whether a model can turn business requirements into functioning software workflows, beyond merely producing a plausible document or clicking the correct button. Its results show progress on that research problem while leaving a substantial gap between partial credit and dependable completion of an entire business process.

Ironclad employees and OpenAI staff familiar with the software helped define the tasks, with individual assignments scored against multiple criteria. The training approach used public contracts and synthetic material, rather than private customer agreements. OpenAI also reported a higher score for an internal development model, which should not be confused with a generally available product. Ironclad technology chief Sunita Verma said useful AI must do more than individual actions while “preserving the controls teams rely on.” That requirement makes the choice of evaluation domain consequential: a contracting process can look complete while still applying the wrong approval route or allowing an exception that the business intended to prohibit.

The percentages require careful interpretation. A criterion score is not a claim that 55% of complete contracts were processed successfully, nor does it reveal how often a particular serious error occurred. A task with many correct elements can still fail its business purpose if one important condition is wrong. As an illustrative example, the quality of an intake form cannot compensate for an approval threshold that routes spending incorrectly. This is a general interpretation of the measurement method, not an additional observed failure in the experiment. Customers evaluating comparable agents would need the distribution of errors and the importance of missed requirements, alongside any average score.

The timing comparison also separates model progress from commercial return. A faster attempt can reduce waiting and computer occupation, but the relevant operating cost includes checking, correction and any repeated attempt. If verification remains expensive, a large improvement in execution time may translate into a much smaller reduction in total handling time. If checking can be made systematic and inexpensive, the same improvement could become commercially useful sooner. Neither outcome has been demonstrated by a customer deployment in this disclosure. The experiment identifies a promising direction for training and evaluation; it does not supply the labour cost, transaction volume or error consequences needed for a financial forecast.

The partnership model gives OpenAI access to detailed definitions of professional work that are difficult to infer from general internet material. Software providers can specify what a successful configuration should do across different cases and how to test it. That creates a potential route to better models while allowing the provider to shape the behaviour expected inside its product. The commercial terms were not disclosed, so the research should not be counted as a new revenue contract. Its strategic value is more indirect: if task definitions and evaluation environments improve reliability, OpenAI may make future agents easier for enterprises to adopt and easier for their suppliers to support.

Analysis

Astra’s score improved by 13.4 percentage points, about 32% relative to Sol, while estimated attempt time fell about 48%. The asymmetry matters: speed improved substantially, but the score remains far from universal correctness. OpenAI’s route to value lies in converting that progress into supervised work whose total cost, including review, beats the existing process. Ironclad contributes the business rules and environment needed to judge that claim. Until the cost of handling residual errors is measured, the results support further development rather than a forecast of equivalent reductions in legal or procurement staffing.