Lilith.
⌕
Editorial illustration: Astra completed just 55% of contract workflows, which is why Ironclad's test matters
Lilith illustration · editorial remix

OpenAI and Ironclad turned complex contracting workflows into 11 research tasks. GPT-6 Astra averaged 55.0%, compared with 41.6% for GPT-5.6 Sol. Estimated average time per attempt fell from 37.0 to 19.2 minutes.

Eleven tasks measure the finished process rather than correct clicks

The tasks span legal, commercial and procurement work. They include setting up a nondisclosure agreement, creating a procurement approval process and updating a reusable contract clause for a selected jurisdiction. OpenAI estimates that an experienced user would need 30 to 40 minutes per task on average.

Each task was scored against 8 to 50 criteria. Ironclad supplied hosted test environments, while researchers created synthetic training tasks and used reinforcement learning to improve the models. An internal development model reached 63.7%, but OpenAI does not tie that result to an available model.

Enterprise software is becoming a training ground for frontier models

The collaboration model matters more than the legal vertical alone. A software vendor is not merely adding an AI feature above its API. It brings domain experts, a test environment and a detailed definition of success directly into frontier model development.

For product teams, this is a method for turning a vague agent request into a measurable task. Completing a form is insufficient. The agent must preserve thresholds, approvals and exceptions across the workflow, then verify that the resulting process behaves correctly under different inputs.

Fifty-five percent is research progress, not permission to work alone

The comparison uses 11 tasks designed by the partners inside Ironclad's hosted environment. Astra ran with Max reasoning and Sol with High reasoning, the settings where each scored best. Without an independent set, result variance and an account of error severity, these figures do not establish production reliability.

The announcement also does not launch a new publicly available Ironclad feature. OpenAI explicitly says human oversight still matters. In contracting, one missing approval branch can outweigh nine correctly completed fields.

Independent evals and error audits will decide whether this can ship

The next useful signal is performance on unseen processes, repeated runs and a changed interface. Enterprises will also need action logs, permission enforcement, uncertainty escalation and a measure of how much human correction remains after the agent's 19.2-minute attempt.

OpenAI is seeking a small number of additional software partners that can provide concrete failures, experts, a secure environment and usable data. If that collaboration produces transferable evals, it may matter more than Astra's score itself.

Lilith's verdict

Astra finished the contract workflow in 19.2 minutes, but left nearly half the criteria scattered by the finish line. The lawyer can take a hand off the mouse, not off the responsibility.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗