
Somebody is about to ask you to commit another year of AI budget. The material you have to justify it with is mostly recycled: the same trend lists, the same unsourced dollar figures, the same confident predictions that nobody ever goes back and checks. That is a poor basis for a decision that size.
So this is a different kind of piece on the future of generative AI. It starts by scoring the predictions made for 2025, including the ones this page itself got wrong. Then it says where agents actually are against a measured benchmark rather than a slogan. Then it names the conditions under which spending more on generative AI is the wrong call.
Every number here carries a date and a source. That matters more than usual on this topic, because a forecast with no date attached cannot be checked, and one that cannot be checked is not evidence.
The future of generative AI is a shift from producing content to completing work. Models that write, draw and summarize are now the base layer. What is being built on top of them are agents that execute multi-step tasks, and, further out, world models that hold an internal representation of how things behave.
That reframing matters because it changes what you are buying. A content generator is a feature. An agent is a process owner, and process owners need supervision, evaluation and an escalation path.
Text and image generation is settled. It works, it is cheap, and it is table stakes in most software you already pay for. If you need the ground floor, our a complete guide to generative AI covers the mechanics. The rest of this article is about the floors above it.
Almost nobody scores their own generative AI predictions, which is why predictions stay cheap to make. Here is the last round, checked against evidence published in 2026.
Two of those rows deserve a note.
The second row is ours. This page predicted in February 2025 that multi-agent systems would be automating banking and medical research workflows by the end of that year. The direction was right and the timeline was not, and the acronym we used for them never became a term anyone else adopted. Stanford HAI’s 2026 AI Index, published in April 2026, is what settles it: agents on OSWorld, a benchmark of real computer tasks, went from 12% task success to about 66%. A system that fails one attempt in three does not run a bank.
The third row needs its caveat stated plainly. Fortune’s August 2025 report on MIT’s NANDA study is the source most people cite for enterprise AI failure. The finding is 13 months old, and its headline framing has been publicly disputed as overstated. We use the 5% success figure rather than the 95% failure figure, and we treat it as contested. Inheriting an overclaim is the exact habit this article is trying to correct.
Three tests will tell you whether an AI prediction is worth anything. Does the claim name a date and a source you can open. Does it name a benchmark, or only a direction. And has its deadline already passed without anyone going back to check.
Apply the last test generously and you will find expired predictions still presented as the future on pages ranking today. One page currently on the first results page for this query still describes the 2026 training-data deadline as an upcoming event.
Now apply all three to this article. The freshest figure here is from August 2026. The newest dataset is from April 2026. We will review this page in March 2027, when the next AI Index lands, and extend the table above.
Five shifts have real evidence behind them. The rest of the usual generative AI trends list is either already ordinary or still speculation.
Agents reach supervised production, not autonomy. This is the center of gravity now, and the number to hold onto is the OSWorld figure above: roughly two thirds success on real computer tasks. Good enough for work a human reviews. Not good enough for work nobody checks. So the future of AI agents is really a question about supervision design rather than raw capability. If you are sorting out where one ends and the other begins, agentic AI versus generative AI draws the line more carefully than most vendor material does.
Multimodal generative AI becomes the default input. Text, image, audio and video in one model stops being a feature and becomes the assumption. The practical effect is on your data: documents you never digitized properly are suddenly readable.
Models get smaller and more specific. The interesting frontier is no longer one enormous general model. It is a domain-tuned model small enough to run where your data already lives. Organizational AI adoption reached 88% according to the AI Index, and adoption at that level pushes cost and control to the front of every architecture conversation.
World models are the honest answer to “what comes next.” This is the least covered and most substantive item on the list. The argument, made most prominently by Yann LeCun, is that the next real advances will not come from scaling language models further. They will come from systems that build an internal model of how the world behaves and can plan against it. Nothing here is production technology yet. It is where the research money is going.
Physical AI and the energy bill both arrive. Generative models are moving into warehouses and robotics. At the same time, compute and power are becoming the binding constraint rather than talent. The US now hosts 5,427 data centers, more than ten times any other country, and that concentration is starting to show up in local power and water debates.
Three more things get named on almost every outlook piece, so here they are honestly and briefly. Hyper-personalization already happened and is no longer a trend. Generative AI in threat detection is real and is now a feature of security products rather than a project. Deepfakes and synthetic misinformation keep getting cheaper, and documented AI incidents rose to 362 in the AI Index count, up from 233 the year before.
Enterprise generative AI adoption is broad and shallow, and the two numbers that show it are worth putting side by side. Organizational AI adoption reached 88%. US population adoption of generative AI sits at 28.3%, which ranks 24th globally, behind Singapore at 61% and the UAE at 64%.
Those measure different populations, so they are not a contradiction. Read together they say something useful: your company has almost certainly adopted AI, and your people mostly have not.
The spending pattern compounds it. MIT’s NANDA study found more than half of generative AI budgets went to sales and marketing tools, while the largest measured returns sat in back-office automation. Money went where the demos were exciting rather than where the work was repetitive.
That finding matches what we see in delivery. One procurement enterprise had finance staff reconciling invoices, purchase orders and goods receipt notes by hand across PDFs, scans and structured records. The AI-powered invoice reconciliation platform we built for them used LLM extraction on AWS Bedrock with automated validation, and it reached production: an 80% reduction in manual reconciliation effort and 2X faster payment processing. Unglamorous, document-heavy, measurable.
The cause of the failures is not model quality. NANDA identified the learning gap, meaning tools that never adapt to a real workflow and organizations that never adapt around them. An independent dataset points the same way: the AI Index found 73% of AI experts expect a positive effect on how people do their jobs, against 23% of the public.
A 50-point belief gap is what generative AI and the future of work actually turn on. Not model choice. Whether the people doing the job think the tool helps them. Change management is the binding constraint on most 2027 plans, and it is the line item that gets cut first.
Be aware of our position here before you read the table. Entrans sells delivery partnership, and the evidence below favors partnering. We think the evidence is good, and you should still weigh it knowing who wrote it.
The build row is not a straw man. Building is right more often than vendors admit, and the three conditions in its last cell are the test. Miss any one of them and the 67% versus one-third comparison starts to apply to you.
Our generative AI consulting work usually starts by arguing a client out of one of these routes rather than into it, and our broader AI engineering and delivery services exist to make the partner row true rather than aspirational.
Governance on this topic is usually a paragraph about bias and a nod to the EU AI Act. That is a value, not a control, and values do not survive an incident.
The measured picture is getting worse, not better. The AI Index reports that responsible AI is not keeping pace with capability: capability benchmarks are reported almost universally while responsible-AI benchmark reporting stays patchy. Documented incidents rose to 362 from 233 the year before. Research also found that improving one responsible-AI dimension, such as safety, can degrade another, such as accuracy. Safety is a tradeoff you budget for, not a box you tick.
Four controls are the practical minimum. Each needs a named owner, an evidence artifact, and a cadence.
On regulation, describe and plan, do not promise. The EU AI Act’s obligations phase in on their own schedule and the NIST AI Risk Management Framework is voluntary guidance rather than a certification. No architecture, tool or vendor makes an organization compliant with either. That judgment belongs to your counsel and your auditor.
There is a commercial reason to take this seriously beyond the legal one. Trust is low and falling: the AI Index found the US reported the lowest confidence of any surveyed country in its own government’s ability to regulate AI, at 31%. Customers who do not trust the rules will look harder at yours.
Four conditions under which more generative AI spending is the wrong call next year.
Capability is jagged, so per-task reliability is unpredictable. Gemini Deep Think earned a gold medal at the International Mathematical Olympiad. The best model reads an analog clock correctly about half the time, at 50.1%. Benchmark wins do not generalize. Better move: pick workloads where a wrong answer is cheap to catch, and stay away from anything whose failure mode is silent.
Nobody owns the learning gap. If no named person owns workflow integration and the feedback loop, the pilot joins the majority that showed no measurable P&L impact. This is the most common reason good technology produces nothing. Better move: fund the change management before the model, and name the owner before you approve the budget.
Your users are less ready than your board thinks. The 88% organizational figure against 28.3% population adoption is the shape of the problem. Better move: measure actual internal usage of what you already bought, honestly, before forecasting a productivity gain from something new.
The investment case is unproven at this scale. Goldman Sachs Research forecast in August 2026 that global AI investment would exceed $1 trillion in 2026, roughly $1.8 trillion cumulatively, rising from 1.8% of US GDP to 2.8% by 2028. Whether that is a bubble is genuinely open, and the same analysis supplies the rebuttal: those levels sit inside the 2% to 5% of GDP peaks seen in earlier general-purpose technology buildouts. We take no view, and none of this is investment advice.
Here is the thing we tell clients that costs us work. Do not fund a copilot for a process nobody has documented. You will spend two quarters discovering that the process was the problem, and the model will get the blame.
The future of generative AI will be decided inside individual workloads, not in forecasts like this one. So if you are setting AI direction for next year, the first useful question is not which model to use. It is which of your workloads can carry one. Entrans runs a generative AI readiness assessment: we map candidate workloads against the four failure conditions above, size the change-management work honestly, and hand back a sequenced roadmap with the workloads that will not pay off marked as such. 150+ AI projects delivered. Talk to an Entrans AI architect and bring the workload your board is most excited about.
Agents come next, then world models. The future of generative AI runs through systems that execute multi-step work rather than produce content, and then through models that hold an internal representation of how the world behaves. Text and image generation becomes the base layer that other systems are built on.
Because the gap is organizational, not technical. MIT’s NANDA initiative reported in August 2025 that roughly 5% of pilots reached rapid revenue acceleration, and identified workflow integration and the learning gap as the cause rather than model quality. Budgets also skewed to sales and marketing while measured returns sat in back-office automation.
No, on current evidence. Stanford HAI’s 2026 AI Index, published April 2026, reports capability accelerating rather than plateauing, with performance on SWE-bench Verified rising from 60% to near 100% in a single year. Gains are uneven across tasks, so faster progress does not mean uniform reliability.
Nobody credible puts a date on it. The honest read from current benchmarks is that capability is jagged: a model earned a gold medal at the International Mathematical Olympiad while the best model reads analog clocks correctly about half the time. Performance that uneven is not a system approaching general intelligence.
The evidence favors document-heavy back-office work over customer-facing content, whatever the industry. Finance, insurance and healthcare administration all run on high volumes of unstructured documents, which is where measured returns have been strongest. Workload shape matters more than industry: repetitive, reviewable, already instrumented.
Unresolved, and the same data supports both readings. Goldman Sachs Research forecast $1 trillion of global AI investment for 2026, rising toward 2.8% of US GDP by 2028. That is a very large number, and it also sits inside the 2% to 5% range seen in earlier general-purpose technology buildouts.


