Quick Takeaways
What you'll learn in this article
- 1
4 scored 75 percent on OSWorld-V this week, quietly crossing the 72
- 2
4 percent human baseline for real-world software workflows
- 3
This is the inflection point enterprise AI has been waiting for, and the shift from copilot to coworker is going to be messier than the hype cycle suggests
Keep reading for detailed implementation, code examples, and real-world results
The Line Nobody Noticed Got Crossed
On Tuesday, in a release that was overshadowed by a Stanford index report, a Broadcom compute deal, and roughly forty smaller news cycles, OpenAI shipped GPT-5.4. The release notes were characteristically understated. A one-million token context window. Improved tool use. A new training recipe for sustained multi-step execution. And buried near the bottom, almost as an afterthought, one number: 75 percent on OSWorld-V.
You could be forgiven for scrolling past it. OSWorld is not a name that circulates outside a narrow slice of AI researchers. It does not have the cultural weight of MMLU or the meme-ready drama of Humanity's Last Exam. It is not the kind of benchmark that produces viral demos or journalist-ready superlatives. And yet this week, for the first time, a frontier model crossed a threshold on OSWorld that matters far more than any of the benchmarks that get headlines: the human baseline.
The human baseline on OSWorld-V is 72.4 percent. GPT-5.4 scored 75. That is not a rounding error. That is not within the noise of evaluation variance. That is a capable model doing real software work, on real applications, across real multi-step tasks, and doing it measurably better than the average human professional that the benchmark was calibrated against.
Every enterprise AI conversation from this point forward has to reckon with that number.
| model | score |
|---|---|
| GPT-4.1 | 38.1 |
| Claude Opus 4.5 | 52.3 |
| GPT-5 (launch) | 58.9 |
| Gemini 3.0 Pro | 61.2 |
| Claude Opus 4.6 | 67.4 |
| GPT-5.4 | 75 |
| Human baseline | 72.4 |
What OSWorld Actually Measures
To understand why this is a bigger deal than another percentage point on MMLU, you have to understand what OSWorld is actually asking the model to do.
OSWorld is a benchmark developed to evaluate computer-use agents inside a live, containerized desktop environment. Not a simulation. Not a toy sandbox. An actual Ubuntu VM running Firefox, LibreOffice, VS Code, Thunderbird, GIMP, the file manager, the terminal, and a clutch of real-world web applications. The model gets a task written in plain English and a screenshot of the desktop. It returns mouse and keyboard actions. It sees the consequences of those actions in the next screenshot. It keeps going until the task is done or it gives up.
The tasks are not "sort this list" or "fill out this form." The tasks look like this:
"Your project manager asked you to export the Q3 sales data from the Salesforce report, pivot it by region in a spreadsheet, generate a chart of year-over-year change, paste the chart into an email draft addressed to the regional VPs, and schedule it to send at 8 AM Monday."
That is one task. It involves three applications, a web service, a file system, a scheduler, and the kind of lateral reasoning required to recover when the Salesforce export comes back with an extra header row or the spreadsheet flags a currency column as text. It is not a single-step capability test. It is a multi-application, multi-hour, error-tolerant knowledge-work scenario. And human professionals recruited to calibrate the benchmark solved 72.4 percent of tasks in this category.
OSWorld-V is the latest revision, hardened against the kinds of memorization and trajectory-level contamination that plagued earlier computer-use benchmarks. It includes tasks written after all current model training cutoffs. It rotates application versions. It uses adversarial paraphrasing to prevent models from pattern-matching tasks they have seen before.
A 75 percent score on OSWorld-V is not a party trick. It is a model that, on average, can take a knowledge-work assignment written by a manager in normal English and produce the correct output more reliably than the human who would have been assigned the same task.
Why This Is Different From Every Prior AI Milestone
For the last three years, the AI industry has been saturated with benchmark-crossing events that did not matter as much as they sounded. GPT-4 passed the bar exam in 2023. Models have been beating human experts on graduate-level physics questions for two years. Humanity's Last Exam was supposedly the final bastion, and then the top labs started scoring above 50 percent on it by early 2026.
These milestones were real, but they were also misleading, because they measured a specific and quite narrow capability: the ability to produce the correct text answer to a closed question. That is a useful skill. It is nothing like the skill of getting actual work done.
Knowledge work has almost nothing in common with exam-taking. It is not a single question with a single answer. It is a fuzzy goal, ambiguously specified, requiring interaction with five applications, with decisions cascading across steps where a mistake in step two makes step five impossible. The correct answer at each micro-step is context-dependent. The tools required are proprietary and idiosyncratic. The feedback loop is slow. The "done" criterion is often social rather than technical โ you are done when your boss is satisfied, which is a moving target.
This is the skill that OSWorld-V measures. And this is the skill that, as of Tuesday, an AI model is demonstrably better at than the median human it is being compared to.
| quarter | examBenchmarks | osworldScore |
|---|---|---|
| Q1 2024 | 91 | 12 |
| Q3 2024 | 95 | 19 |
| Q1 2025 | 97 | 28 |
| Q3 2025 | 98 | 44 |
| Q1 2026 | 99 | 67 |
| Q2 2026 | 99 | 75 |
The gray line is exam-style benchmarks, which models have been saturating for a while. The red line is OSWorld, which measures actual software work. Until this quarter, the gap between "passes the exam" and "can do the job" was enormous. This week the red line finally crossed the one that matters.
The Copilot Metaphor Is Officially Dead
For the last two years, every enterprise AI vendor pitched their product as a copilot. The metaphor was deliberate. A copilot is a skilled partner who assists but does not replace the pilot. A copilot handles the radio and the checklists and the weather updates, but the pilot makes the decisions. Responsibility stays with the human.
The copilot metaphor was useful. It let vendors sell powerful AI tools to enterprises that were terrified of taking humans out of the loop. It let procurement teams tick the compliance boxes. It let HR teams reassure employees that their jobs were not at risk, just augmented. It gave the entire industry a soft-landing narrative for the displacement that was obviously coming.
That metaphor is now dead. A model that scores above the human baseline on a real-world software benchmark cannot honestly be sold as something that assists the human. It is a functional peer. It is a coworker, in the strict sense that the word implies: a separate agent that shares work on roughly equal footing.
The replacement metaphor the industry is already reaching for is autonomous coworker, and you will see it in every enterprise AI pitch deck within the next six weeks. OpenAI's quiet Tuesday announcement that they are "shifting focus to business users" with a new professional-grade model is the first of many announcements that will lean into this framing. Anthropic has been building toward it for a year with Claude Cowork, which is probably why OpenAI's pivot looks so urgent.
The framing matters. A copilot gets sold to the individual worker. An autonomous coworker gets sold to the budget owner who signs the worker's paycheck. These are different buyers with different incentives, and the economics of the second buyer are roughly an order of magnitude larger than the first.
| category | annualSpendUSD |
|---|---|
| Individual Seat (Copilot) | 240 |
| Team Seat (Copilot) | 900 |
| Autonomous Coworker (per unit) | 18000 |
| Displaced FTE Salary | 95000 |
This chart is why the whole enterprise AI industry is about to pivot overnight. The average annualized spend for a "copilot" license is in the low hundreds of dollars per seat. The pricing emerging for autonomous-coworker products is in the tens of thousands of dollars per unit, because the buyer is explicitly comparing the cost against a displaced full-time employee salary, not a productivity tool.
A 75-percent OSWorld-V model at $18,000 per year is not competing with Microsoft Copilot. It is competing with a $95,000-per-year knowledge worker. That is a 5x margin on unit economics. And that is exactly the kind of unit economics that lets a company like OpenAI justify the capital intensity of frontier training runs.
Why OpenAI's Enterprise Pivot Is Suddenly Urgent
The timing of OpenAI's announcement this week is not coincidence. For the last eighteen months, Anthropic has been quietly building the better enterprise story. Claude Cowork, which triggered the SaaSpocalypse in February, was the first product explicitly designed around the autonomous-coworker model rather than the copilot model. It priced like coworker software. It sold to procurement. It talked in the language of full-time-equivalent displacement.
OpenAI's enterprise story, by contrast, was anchored to ChatGPT Team and ChatGPT Enterprise โ products that looked and priced like copilots, bolted on top of a consumer experience. That strategy worked beautifully in 2024 and 2025 when the product-market fit was "chatbot for workers." It stopped working the moment the benchmarks said the models could do the work themselves.
GPT-5.4 and the OSWorld-V score give OpenAI the ammunition they need to reposition. The Tuesday release notes and the leaked "Spud" codename for the next professional-grade model are, when you read between the lines, an explicit repositioning away from the per-seat consumer-product business and toward the per-unit autonomous-worker business that Anthropic has been building.
The competitive pressure is real. Anthropic's annualized revenue is reportedly approaching $19 billion. OpenAI's is above $25 billion, but most of that is still consumer. Anthropic's revenue mix is heavily skewed toward enterprise, which means their dollar-per-user yield is meaningfully higher and their unit economics are more defensible as frontier training runs continue to get more expensive.
| company | revenueBillionsUSD |
|---|---|
| OpenAI Consumer | 16.2 |
| OpenAI Enterprise | 8.9 |
| Anthropic Consumer | 4.1 |
| Anthropic Enterprise | 14.7 |
OpenAI has more total revenue, but Anthropic has more enterprise revenue. If the enterprise opportunity is an order of magnitude larger than the consumer opportunity (which the unit economics above suggest it is), then Anthropic is currently winning the race that actually matters, and OpenAI's pivot is an acknowledgment of that.
The Infrastructure That Makes Autonomous Coworkers Possible
The OSWorld-V score is not just a model-quality number. It is also a proxy for an infrastructure story. An autonomous coworker that can do multi-application software work at human-baseline reliability requires three things beyond raw model intelligence: long context, reliable tool access, and the compute to run inference at enterprise scale.
The first is solved. GPT-5.4 ships with a one-million token context window by default. Claude Opus 4.7 has similar context scale. Gemini 3.1 Pro has been at two million for months. A knowledge-work task that spans a dozen documents, a dozen emails, and a dozen Slack threads is well inside the working memory of every frontier model released this year.
The second is solved by the Model Context Protocol, which crossed 97 million installs in March. MCP has quietly become the USB-C of agent tool use. Every major enterprise software vendor now ships an MCP server. Salesforce, ServiceNow, Workday, SAP, Microsoft 365, Google Workspace, Atlassian, HubSpot, Snowflake, Databricks, and the long tail of SaaS vendors below them all support MCP connections as a default deployment option. For the first time in the history of enterprise software, an AI agent can reach into the actual systems of record without bespoke integration work for every customer.
The third โ the compute โ is the one the industry is still scaling into, and it is the reason the Anthropic / Google / Broadcom deal for 3.5 gigawatts of next-generation compute was the real infrastructure story of this month. An autonomous coworker that costs $18,000 per year and works 40 hours per week and is backed by a frontier model burns a non-trivial amount of inference compute. Serving millions of them simultaneously requires roughly an order of magnitude more infrastructure than the current copilot products consume. The buildout is happening, but the capacity constraint is real, and it is the single biggest reason you should not expect the full autonomous-coworker transition to happen in the next twelve months.
| year | inferenceGW | trainingGW |
|---|---|---|
| 2024 | 2.1 | 4.8 |
| 2025 | 6.4 | 9.2 |
| 2026 | 18.7 | 17.5 |
| 2027 (projected) | 52.1 | 28.4 |
| 2028 (projected) | 121.3 | 41.8 |
The inference curve has officially diverged from the training curve. For most of the last five years, compute growth was dominated by training larger and larger models. Starting this year, inference compute for production agent workloads is eclipsing training. This is the signature of an industry that is deploying, not experimenting.
What Actually Happens in the Enterprise Now
The temptation when a milestone like this gets crossed is to project straight-line exponential change. AI Twitter will spend the next two weeks proclaiming the imminent end of knowledge work. The mainstream press will run a few "robots are coming for your job" cycles. Some consultancy will publish a report claiming 40 percent of white-collar work will be displaced by 2028. These projections are wrong in the details even when they are directionally correct.
What actually happens is slower, lumpier, and more institutionally complicated than the hype suggests. Here is the pattern you should expect based on how previous enterprise-tech transitions played out, refined for what OSWorld-V parity specifically unlocks.
Months 1 to 3 (now through July 2026): The frontier labs and a handful of enterprise-AI-native startups ship "autonomous coworker" products that explicitly claim human-baseline-or-better performance on specific high-value job functions. The initial targets will be SDR work, L1 customer support, financial analyst research, paralegal document review, and specific categories of software development. The pricing will be per-unit per-year, not per-seat. The sales motion will go directly to budget owners, bypassing IT procurement when possible.
Months 4 to 9 (August 2026 through February 2027): The first big enterprises publish case studies claiming 30 to 50 percent FTE reduction in specific job functions. The studies will be technically accurate but strategically misleading โ the "displaced" headcount will mostly be natural attrition that is not backfilled, not mass layoffs. This phase generates the worst press the AI industry has yet seen, because the cumulative attrition across the economy starts to show up in labor market data.
Months 10 to 18 (March 2027 through September 2027): The legal and regulatory response begins in earnest. The first wrongful-termination suits against companies that displaced human workers with agents that have demonstrable failure modes start to produce rulings. Several US states introduce "AI worker disclosure" legislation. The EU extends the AI Act with specific provisions for autonomous-worker deployments. The industry enters a compliance-heavy phase that slows deployment velocity without reversing it.
Months 19 to 30 (October 2027 through September 2028): The stabilization period. Enterprises that deployed autonomous coworkers well pull ahead on margin. Enterprises that deployed badly produce high-profile failures that scare the middle of the market. A durable equilibrium starts to emerge where autonomous coworkers are a standard part of the workforce mix, typically at 20 to 40 percent of headcount in their target job functions. This is the new normal.
| phase | ftesDisplacedThousands |
|---|---|
| Months 1-3 | 85 |
| Months 4-9 | 420 |
| Months 10-18 | 780 |
| Months 19-30 | 1450 |
These displacement numbers are my estimates, grounded in the OSWorld-V capability benchmark, the compute capacity actually coming online, the enterprise procurement cycles I have observed in the Anthropic rollout over the last three quarters, and the historical base rate of enterprise tech adoption. They are likely to be wrong in the specifics and right in the magnitude. Cumulatively, roughly 2.7 million US knowledge-work FTE positions get displaced or significantly transformed in the thirty months following OSWorld-V parity.
That is a non-trivial number, but it is also substantially smaller than the apocalyptic projections you will see in the next news cycle. The labor market absorbs roughly 150,000 to 250,000 positions of structural change per month in normal conditions. An AI-driven displacement rate of roughly 90,000 FTE per month over two and a half years is absorbable, painful in specific sectors, and likely to accelerate the structural re-skilling investments that have been underfunded for fifteen years.
The Three Things Enterprise Buyers Should Actually Do
If you are a technology leader at an enterprise reading this and trying to figure out what it means for your 2027 planning cycle, the answer is surprisingly concrete. There are three specific actions worth taking in the next 90 days.
First, audit your job-function taxonomy against OSWorld-style capability benchmarks. Not every knowledge-work role is equally exposed. The roles most exposed to autonomous-coworker displacement are the ones that look most like OSWorld tasks: bounded, software-mediated, with clear done- criteria and recoverable error modes. Roles that require physical presence, sustained relationship-building, high-stakes judgment on ambiguous values questions, or creative originality are substantially less exposed. Your org chart probably does not reflect this taxonomy yet. Build it.
Second, run three to five focused pilots of autonomous coworker products on specific job functions, with explicit success criteria tied to unit economics, not "productivity." Copilot pilots failed to produce durable enterprise value because they tried to measure something intangible (productivity) across a population of users who were incentivized to report improvement whether or not it existed. Autonomous coworker pilots should be measured against FTE-equivalent output with specific accuracy and reliability SLAs. If the pilot cannot produce output equivalent to an FTE at the target price point, the product is not actually an autonomous coworker yet, regardless of what the vendor says. Your pilots will be most honest if they compare directly against the cost and output of existing employees doing the same work.
Third, invest heavily in the supervisory layer. The biggest mistake most enterprises will make in the next two years is treating autonomous coworkers as drop-in FTE replacements. They are not. They are agents that produce output at FTE-equivalent or better rates, but with failure modes that are different from and less legible than human failure modes. They need supervision. The supervisory function is itself a job โ probably a senior, well-compensated job โ and the enterprises that build a supervisory discipline will capture more durable value than the ones that deploy agents without one. This is the same pattern that played out with orchestration layers for multi-model agents: the real enterprise value is in the management, not the worker.
What This Means for Anthropic, Google, and the Also-Rans
A model of the competitive dynamics after OSWorld-V parity looks different from the competitive dynamics that existed last quarter. The frame has shifted from "who has the smartest model" to "who ships the most reliable autonomous coworker at the lowest unit economics."
Anthropic is best positioned in the short term. Claude Cowork has the product-market fit. Claude Opus 4.7 is at or near the capability frontier. The 3.5 GW Broadcom / Google compute deal closes the capacity gap that was their biggest weakness. The enterprise sales motion is already built. The brand in enterprise AI is stronger than anyone else's. If the market moves the way this analysis suggests, Anthropic's revenue trajectory through 2027 probably justifies a valuation at or above OpenAI's.
OpenAI is pivoting fast and has the largest distribution, the largest war chest, and the strongest research bench. The Tuesday announcement of a business-focused model is the right move. The challenge is that their product DNA is consumer, and autonomous-coworker selling is enterprise. They will need to build or acquire an enterprise go-to-market motion roughly on par with Anthropic's in the next nine months, or they cede a generational market. It is an open question whether a company this large can pivot that fast.
Google is the dark horse. Gemini 3.1 Pro is a competitive model. Google Workspace has the largest enterprise knowledge-work footprint in the world outside of Microsoft 365. The existing customer relationships, the existing context about what users actually do, and the in-house compute capacity are all advantages. The challenge is Google's historical difficulty converting enterprise technical advantages into enterprise revenue. If they execute, they win. The base rate on Google executing on enterprise software is mixed.
Microsoft is in an interesting position because they are positioned as both an AI provider (via the OpenAI partnership) and the default platform for enterprise knowledge work (via Microsoft 365 and Azure). An autonomous-coworker market that matures on top of Microsoft's existing distribution accrues a lot of value to Microsoft, almost regardless of which model provider wins. The biggest strategic risk to Microsoft is that the autonomous coworker market dis-intermediates the Microsoft 365 productivity bundle. If the coworker works inside Salesforce and ServiceNow directly via MCP, without ever touching Word or Outlook, the bundle loses its gravity.
Meta, Amazon, xAI, and the long tail of labs have specific bets that may or may not matter. Meta's open-source play (Llama 5 and the surrounding ecosystem) is genuinely disruptive in the small-model and edge-deployment segments, but not yet at the frontier capability level required for OSWorld-V parity. Amazon's Bedrock strategy makes them a critical enterprise AI infrastructure provider without owning the capability frontier. xAI is an outlier bet on a specific user base. None of these are yet playing in the autonomous-coworker category at OSWorld-V parity, though any of them could be within eighteen months.
| company | positioning | distribution | compute | productFit |
|---|---|---|---|---|
| Anthropic | 85 | 55 | 70 | 90 |
| OpenAI | 60 | 95 | 80 | 55 |
| 65 | 85 | 90 | 45 | |
| Microsoft | 70 | 95 | 85 | 60 |
These are rough scores across the four dimensions that actually matter for autonomous-coworker market share: positioning (does the company talk about their product the way a buyer wants to hear about autonomous workers), distribution (does the company have existing enterprise relationships to sell through), compute (does the company have the infrastructure to serve the workload), and product fit (does the shipping product actually work like an autonomous coworker today). The company with the most durable position is usually the one with the most balanced profile, not the one with the single highest peak.
But Is OSWorld-V Being Gamed?
The reasonable skeptical response at this point is: benchmarks get gamed, and we have watched models saturate benchmarks that turned out not to translate into real-world capability over and over. Why is OSWorld-V different? This objection deserves a real answer rather than a dismissal, because it is the right question to ask.
Three things make OSWorld-V substantially harder to game than most prior benchmarks.
The first is that the evaluation happens inside a live, stateful environment rather than against a static dataset. A model cannot memorize the correct trajectory, because the environment responds differently each time the task runs. A file system changes. A network call might be slow. An application might auto-update its UI between evaluation runs. The model has to actually interact with a real system, which is categorically different from producing text that matches a reference answer. You cannot pre-train your way to a high OSWorld-V score the way you can pre-train your way to a high MMLU score.
The second is that the task distribution is held out in ways that are verifiably not in training data. The OSWorld-V maintainers publish the task construction methodology but not the specific tasks used in the held-out evaluation set. New tasks are added quarterly. The application versions used in the VM are updated to track current releases. The adversarial paraphrasing of task instructions makes pattern-matching against training-data task descriptions ineffective. None of this is a perfect guarantee against contamination, but it is meaningfully more rigorous than the benchmarks that came before.
The third, and most important, is that OSWorld-V scores correlate tightly with real-world deployment outcomes. This is the empirical test that matters, and it is the test that earlier benchmarks failed. A model that scored 90 on the bar exam did not become a functioning lawyer overnight. A model that scored 99 on MMLU did not become a reliable domain expert. But the internal evaluation data that several frontier labs have shared under NDA โ and that some enterprise customers have begun replicating on their own workloads โ suggests that OSWorld-V scores are genuinely predictive of task-completion rates in production agent deployments. A 10-point OSWorld-V improvement has historically produced roughly proportional improvements in real enterprise pilot outcomes. That correlation is not ironclad, and it will weaken somewhat as the benchmark saturates, but it is substantially stronger than what we had with prior benchmarks.
None of this means OSWorld-V is a perfect measure. It means the benchmark is the best proxy for autonomous-coworker capability that currently exists, and the 75-percent threshold is the most credible capability milestone we have seen.
The Political and Regulatory Dimensions Nobody Is Yet Modeling
There is a version of this analysis that stays strictly inside the economics of enterprise software, and it would still be a legitimate framing. But OSWorld-V parity is not just an enterprise software story. It is a political story, and the political dimensions are about to shape the economic trajectory in ways that are under-modeled in most current forecasts.
Three dimensions matter.
Labor market impact and the political response. Knowledge work has not been meaningfully displaced at scale before. Manufacturing displacement drove decades of political realignment in the United States, the UK, Germany, and Japan. White-collar displacement is likely to drive a faster and more volatile political realignment because the displaced population is educated, politically engaged, and concentrated in cities where political influence is disproportionate. Expect specific policy proposals for retraining subsidies, wage insurance, sectoral pauses, and in extreme cases, outright moratoria on autonomous-coworker deployments in specific sectors.
Liability and accountability frameworks. Enterprise software has historically been licensed on "as is, use at your own risk" terms. Autonomous coworkers cannot be licensed that way, because the failure modes include financial harm, regulatory violations, and reputational damage that scale beyond what traditional software vendors can indemnify against. New frameworks will emerge โ likely through insurance first, then through specialized compliance vendors, then through legislation. The total addressable market for this meta-layer is probably $50 billion per year by 2029.
Geopolitical competition. The US-China AI dynamic shifts meaningfully once autonomous coworkers are deployable. The recent cooperation between OpenAI, Anthropic, and Google to prevent Chinese model copying was not about protecting IP in the abstract โ it was about protecting the capability gap that justifies the premium pricing of autonomous-coworker products. If that gap closes, the entire unit economics of the US frontier labs gets squeezed. Expect increasingly aggressive export controls, increasingly opaque model access policies, and increasingly coordinated industry-government positioning on which markets get which models at which capability tiers.
| Name | Value |
|---|---|
| Enterprise AI Products | 180 |
| Supervisory/Compliance Layer | 50 |
| AI Liability Insurance | 35 |
| Retraining / Reskilling | 28 |
| Regulatory Tech / Audit | 22 |
The market map for the next three years. Roughly $315 billion in annual revenue across these five categories by 2029. The core enterprise AI product category is still the largest, but the surrounding ecosystem is what captures the majority of the economic surplus from the autonomous-coworker transition. The companies that own the supervisory and compliance meta-layer are likely to be the durably most profitable.
The Question That Matters For You
There is a version of this article that ends with an exhortation to "prepare yourself for the future of work." That is the version most AI trend pieces write. It is not a useful ending, because "prepare yourself" is not actually advice. It is vibes.
Here is a more useful ending. The question that matters for you, specifically, as someone reading this article in April 2026, is this:
Is the work you currently do more similar to the kinds of tasks on OSWorld-V, or less similar?
If the work you do is bounded, software-mediated, with clear done-criteria and mostly-recoverable failure modes โ if your workday looks mostly like moving information between applications, applying judgment at specific decision points, and producing outputs that get reviewed by someone else โ then you are in the exposed category. OSWorld-V at 75 percent means there is a credible autonomous coworker on the horizon that can do your job at FTE-equivalent output. You have somewhere between 18 and 36 months to figure out whether your value to your employer is actually the execution work (which is going to be commodified), or the supervisory, strategic, and relational work that sits on top of it.
If the work you do involves physical presence, sustained high-stakes judgment on ambiguous values questions, deep relationship-building with counterparties who expect human-to-human interaction, or sustained creative originality in domains where novelty is the whole product โ then you are in the less-exposed category. Your timelines are longer. The substitution is less direct. The transition is still going to be disruptive, because the economy that surrounds you is going to change enormously in ways that feed back into your work, but you are not the first target.
Most knowledge workers reading this are somewhere between those two categories. That is fine. The useful frame is not "am I going to be replaced yes or no" but "what fraction of my current work is exposed, and what is the supervisory or strategic layer that sits on top of it." That layer is where the durable value concentrates.
The Line That Changed This Week
One more time, because it bears repeating: GPT-5.4 scored 75 percent on OSWorld-V this week. The human baseline on OSWorld-V is 72.4 percent.
That is the line. It got crossed quietly. It is not the AGI line. It is not the singularity line. It is not the end of knowledge work. But it is the line where the word "copilot" stops describing what is being sold, and the word "coworker" becomes honest. It is the line where the enterprise AI market stops being a productivity-tool market and starts being a labor market. It is the line where the economics of frontier AI finally connects to the economics of the human workforce, at the unit level, for the first time.
Everything downstream of that line, from the OpenAI enterprise pivot to the Broadcom compute deal to the regulatory posture of the next two administrations to the quarterly earnings calls of every Fortune 500 CFO for the next three years, makes more sense once you know the line was crossed. Most of the people making decisions about all of these things do not yet know. The ones who figure it out first win the decade.
The compute is online. The models cleared the benchmark. The business models are in flight. The only remaining variable is how fast the institutional and regulatory environment catches up to the capability curve. The bet I would take is that it will catch up slower than most observers expect, because institutions always do, and the transition window is therefore longer than the hype but shorter than the denial. Eighteen to thirty months to the new normal. Not three years. Not five years. Not one.
OSWorld-V at 75. Mark the date.

