Models Will Get Cheaper, Problems Will Get More Valuable: From Today’s AI Boundaries to the Problem-Solving Industry 20 Years From Now

Models Will Get Cheaper, Problems Will Get More Valuable

From Today’s AI Boundaries to the Problem-Solving Industry 20 Years From Now

This article starts from a long-term view: as models, compute, and general-purpose AI become infrastructure, models themselves will get cheaper. What becomes scarce will be problems worth solving, unique data and feedback, the ability to define problems, and the ability to turn them into verifiable results. Starting from the technical boundaries already visible today, we ask whether the AI industry may evolve over the next 20 years from selling models to trading problems and outcomes.

Xiang ZhangChief AI Scientist, China Premium Smart Appliances Innovation Center
1
Early Contributor to and Explorer of Foundational AI Technologies

His PhD advisor was Yann LeCun, recipient of the 2019 Turing Award. He has long worked on foundational AI research and frontier exploration.

2
Extensive Experience in Industrial AI R&D and Delivery

He has end-to-end experience from algorithm R&D to industrial deployment, with projects that materially increased business revenue. His past experience includes Element AI and Google.

3
Two Co-founding Experiences in North American AI Startups

He has twice co-founded AI startups in North America and is preparing to begin a new entrepreneurial venture.

4
Selected for National- and Provincial-Level Talent Programs

Selected for a national-level talent program and the Shandong Taishan Industry Leading Talent program; his academic papers have received nearly 20,000 citations.

The central thesis is:Models will get cheaper, while problems will become more valuable. To understand the AI industry 20 years from now, we cannot simply extrapolate today’s large-model product line. AI is powerful, but today’s dominant form of AI is not all of AI. AI has bubbles, yet over the past two decades it has repeatedly opened a new problem space as the previous growth curve approached its limits. In the long run, the scarce resource may not be models, but problems worth solving—and the ability to organize problems, solutions, and outcomes into transactions.

From Today’s AI Boundaries to a Future Problem-Solving Industry Filter the Noise Media, capital, andnarrative bias Take a Rational Viewof the “Bubble” Narrative Financial cycles ≠technological endpoints Identify theBoundaries What Transformers do well—and what they do not Project theDirection From generating contentto modeling problems directly
Rather than starting from optimism or pessimism, this article projects the future through four layers: information noise, bubble narratives, technical boundaries, and industry direction.
01 · Filter the Noise

Filter the Narrative Before Discussing AI

To discuss the AI industry 20 years from now, we first need to separate AI itself from AI amplified by media, capital, and national-competition narratives. Without first correcting for source bias, it is easy to mistake a benchmark gain for an industrial revolution, a demo for scalable deployment, or a geopolitical judgment for a technical one.

Chinese AI Discourse: Headlines More Easily Amplify Technical Progress

A common pattern in Chinese AI media, especially self-media, is that the technical facts may not be wrong, but headlines and secondary sharing extrapolate local progress into conclusions about the entire industry. Distribution rewards words such as “stunning,” “disruptive,” and “surpassing across the board,” rather than test conditions, failure cases, and boundaries of applicability.

AI Era · Aug 21, 2025 Coverage of DeepSeek-V3.1 used strong phrases such as “stunning release,” “crushing Claude 4,” and “opening a new era of agents.”

What needs filtering is the leap in level: leading on specific benchmarks is not the same claim as entering a “new era of agents.” The headline turns a local technical advantage into a judgment about an entire industry phase.

Synced Pro · May 29, 2025 A minor DeepSeek-R1 update was summarized in the headline as having “upended the large-model landscape.”

Yet the article’s own tests also recorded simple syntax errors and “overthinking.” This shows how an industry-level headline can be more aggressive than the evidence in the body.

How to read the Chinese side: Read the test conditions, failure cases, and scope of applicability first; only then consider claims such as “surpass,” “disrupt,” or “new era.”

U.S. AI Discourse: More Technical Detail, but Chinese AI Is Often Framed Geopolitically

Established U.S. technology media and technical communities usually pay more attention to methods, cost, benchmark conditions, and reproducibility. But when Chinese AI is involved, the framing more readily shifts to “catching up,” imitation, security, military use, and national competition. The technical detail can be useful, but technical facts should be read separately from geopolitical narratives.

The Wall Street Journal · 2026-08-21 When covering Chinese AI talent and startups, the article used “strategic mimicry” to explain part of the progress and emphasized controversial methods such as distillation.

The article also acknowledged that Chinese teams developed efficient techniques under compute constraints, but the overall narrative still placed indigenous technical progress and “imitating the U.S.” within the same competitive frame.

Reuters · 2025-10-27 / 2025-07-09 Reuters coverage of DeepSeek used “robot dogs and AI drone swarms” and “ideological bias” as central angles.

The reporting was well sourced, but it chose to place Chinese AI first in the context of military use, ideology, and national security, so technical capability was often wrapped in geopolitical framing.

How to read the U.S. side: Take the technical detail seriously, but separate technical facts from U.S.–China competition and security narratives.
02 · The AI Bubble

First Understand What a “Bubble” Is, Then Ask Whether AI Will Burst

A “bubble” is first a financial and valuation concept, not a technical one. Around 1999–2000, capital markets focused on dot-com cash burn, profitability, and unsustainable valuations. At almost the same time, however, e-commerce, search advertising, content delivery, and network infrastructure were rapidly maturing.

1999–2002: Financial Repricing and Internet Progress Happened at the Same Time What Financial Markets Saw What the Technology IndustryWas Building • High valuations, cash burn, tighter financing • Many internet companies failed • A major Nasdaq drawdown • “The dot-com bubble burst” became the dominant narrative • Amazon Marketplace / e-commerce model advanced • Google AdWords / Images / News launched • Akamai rapidly expanded CDN infrastructure • Search, advertising, and network services matured At the same time
A financial bubble bursting is not the same as a technology industry failing.

Pessimistic Financial Narratives Contrasted Sharply with Later Technical and Business Progress

Amazon: “Amazon.bomb” in 1999 vs. a Business Breakthrough in 2001

Barron's treated Amazon as a symbol of internet overvaluation; in the same year, seven of eight Wall Street strategists surveyed by Forbes named Amazon among the most overvalued stocks. Yet during the crash Amazon kept expanding Marketplace, reached $3.12 billion in annual sales in 2001, and posted its first quarterly GAAP net profit in Q4 2001.

The Internet Sector: Cash-Burn Anxiety vs. Google’s Business Model Taking Shape

In 2000, Barron's “Burning Up” examined 207 internet companies and estimated that at least 51 could run out of cash within a year. The Washington Post also cited a Merrill Lynch analyst who believed as many as 75% of internet companies might never become profitable and could disappear. Meanwhile, Google launched AdWords, Google Images, and Google News from 2000 to 2002, rapidly building the product and business foundations of search.

Akamai: Listed as a Cash-Burn Risk While CDN Infrastructure Expanded

Akamai was also included in Barron's cash-burn ranking. Yet in the same year, its server count grew from about 2,000 to more than 8,000, network coverage expanded from 100 networks to 473, and annual revenue rose from $4 million in 1999 to $89.8 million. While the financial bubble deflated, internet infrastructure itself was getting much stronger.

Two things must therefore be separated: “Were internet stocks overvalued?” is a financial question. “Were internet technology and business still advancing?” is an industry question. The first mattered greatly in 2000, but it did not imply that the internet had no future.

The 1999–2000 Market Crash Also Closely Coincided with Fed Tightening

Beginning in 1999, the Federal Reserve repeatedly raised the federal funds rate, taking it to 6.5% in 2000. The Nasdaq fell sharply after peaking in March 2000; the broader stock market also declined, although technology and internet shares fell much further.

1999–2001: Financial Tightening and Market Repricing 1999 The Fed begins sustained tightening 2000 MAY Federal funds rate reaches 6.5% 2000 MAR → 2001 Nasdaq falls sharply from its peak BROADER MARKET The broader market is repriced as well

The argument here is: The immediate backdrop to the so-called dot-com bubble bursting was first and foremost a financial cycle. Internet-company overvaluation was real, but the market crash coincided with sustained Fed rate hikes, tighter liquidity, and broad asset repricing, with effects extending far beyond the internet sector itself.

The argument here is also that later U.S. financial and media narratives compressed this complex history: macro tightening, broad market repricing, capital withdrawal, and company failures were gradually condensed into the single label “the dot-com bubble burst.” That label shaped today’s default intuition about “technology bubbles,” while downplaying the simultaneous advances in internet infrastructure, products, and business models.

What Makes AI Different: The Next Curve Rises Before the Previous Bubble Can Sink the Industry

Technologies often follow a path of breakthrough → inflated expectations → bubble → crash → maturity. AI has behaved differently over the past two decades: whenever one technical wave approached a bottleneck, the next advance was already emerging and pushing the industry into a new problem space.

Twenty Years of AI Technology Waves Linear Classification / Regression Deep Learning AlexNet / CNN AlphaGo / RL GPT / ChatGPT Sora / Seedance World Models / Embodied AI 2006 2012 2016 2020 2023 2025+
Individual technologies can overheat, cool down, or collapse, but AI’s overall growth curve has been repeatedly lifted by new tasks and new paradigms.
2006DEEP LEARNING

Traditional Classification and Regression Saturated; Deep Learning Reopened Representation Learning

Machine learning had relied heavily on linear classification, regression, SVMs, and hand-designed features. In 2006, Geoffrey Hinton and Ruslan Salakhutdinov published work in Science showing that multilayer neural networks could learn more effective low-dimensional representations.

Why it carried the next wave: The shift from “humans design features, machines classify” to “machines learn representations themselves.”

2012ALEXNET / CNN

Computer Vision Hit a Traditional-Method Bottleneck; AlexNet Proved the Power of Deep CNNs

AlexNet sharply reduced error rates on large-scale ImageNet recognition and used GPUs to train a large convolutional network, showing that deep learning was not just an academic idea but could decisively outperform earlier methods on real vision tasks.

Why it carried the next wave: Deep learning moved from “trainable” to “clearly stronger,” accelerating the maturation of the vision industry.

2016ALPHAGO / RL

As Vision Commercialized, AlphaGo Moved AI from Perception to Decision-Making

AlphaGo combined deep neural networks, reinforcement learning, self-play, and Monte Carlo tree search. The public imagination of AI expanded from recognizing images to whether machines could make complex decisions through trial and error and long-horizon planning.

Why it carried the next wave: AI’s core problem expanded from perception to decision-making.

Around 2020GPT → CHATGPT

Real-World Reinforcement Learning Hit the Cost of Trial and Error; Generative Models Took Over

After GPT-2, GPT-3 in 2020 further showed that scaling autoregressive language models could produce few-shot capabilities without retraining for every task. In 2022, ChatGPT packaged this capability into a natural-language product anyone could use.

Why it carried the next wave: AI began to acquire general capabilities from the language, knowledge, and code already present on the internet.

After 2023SORA / SEEDANCE

As Chatbots Became Homogeneous, Multimodal Generation Opened a New Experience Frontier

Once text chat became a standard product form, the next growth wave came from images, audio, and especially video. In 2024, Sora demonstrated minute-long high-quality video generation and described video models as “world simulators”; in 2025, Seedance 1.0 further advanced multi-shot generation, consistency, and efficiency.

Why it carried the next wave: The output expanded from “text answers” to complete visual and temporal content.

After 2025WORLD MODELS

After the Robotics Boom Hit a Real-World Data Bottleneck, World Models Gained Momentum

Robots cannot easily learn “how actions change the world” from internet text and video alone. World models began entering mainstream AI narratives: Genie 3 can generate interactive environments, while V-JEPA 2 combines large-scale video pretraining with a small amount of robot interaction data for understanding, prediction, and planning.

Why it carried the next wave: AI’s target expands from “information” on the internet to “state–action–outcome” in the real world.

Conclusion: AI has many bubbles, but AI itself is not a single bubble that will burst as a whole. Individual directions can overheat, cool down, or disappear; yet over the past two decades, whenever one paradigm approached a bottleneck, another had already pushed the industry into a new problem space.
03 · Technical Boundaries and the Next Curve

Why Today’s AI Is Powerful—and Why It Is Not the Endpoint

Start with the Products We Know: What Is Behind ChatGPT, Claude, Gemini, DeepSeek, and Qwen?

The “AI” most people encounter is not an abstract algorithm but a productized large-model service. One layer below the product is the large language model; another layer down, most mainstream LLMs are built around Transformers and their variants.

From Product Layer to the Transformer Core Product Layer ChatGPT · Claude · Gemini DeepSeek · Qwen Model Layer Large Language Models Core Architecture Transformer
To discuss the boundaries of mainstream AI today is first to discuss the boundaries of the Transformer-centered large-model paradigm.

Transformers Succeeded Because They Match Internet Data Extremely Well

“Transformer + large-scale data + tokenization + sequential generation” is exceptionally well suited to text and can also convert internet modalities such as images, video, audio, and code into sequences. This is what gave today’s large models their remarkable generality. But the same structure also determines which data, representations, and solution methods are most natural to them. The key to the next AI curve is not to say simply that “Transformers do not work,” but to distinguish which capabilities can keep improving through scale and which problems are structurally mismatched with the current paradigm.

Three Boundaries of Transformers 1 Boundary 1: Sequential Generation Autoregressive generation must proceed token by token;coupled parallel decisions are forced into a sequence. 2 Boundary 2: Internet Data Large amounts of production, business, and equipment data have never appeared on the internet,leaving a natural gap in the training distribution. 3 Boundary 3: Representing the World Token representations are good at describing symbols, but weaken continuous values,state changes, and the consequences of actions.
1

Sequential Generation: Being Good at “Writing an Answer” Is Not the Same as Solving a System Jointly

The most successful operating mode of mainstream large models is to organize inputs and outputs as token sequences and autoregressively predict the next token. This fits language naturally: sentences unfold word by word, and code also has explicit symbolic order. As models get larger and contexts longer, they become better at preserving information, plans, and logic across long sequences.

But many real-world problems are not about “what to write next,” but about what values hundreds of variables should take at the same time. In a machine, for example, valves, frequencies, temperatures, and pressures may be tightly coupled: changing one variable immediately changes the optimal values of the others. Supply chains, portfolios, and robot joint control have similar structures. Forcing a problem that should be solved jointly into a sequence introduces an artificial order: the model must decide variable A first, then B, C, and D with A already fixed, while the true optimum may require all of them to move together.

Long context therefore mainly answers “how far can the sequence extend?” rather than “can a high-dimensional coupled system be solved jointly?” That is why this article treats non-sequential, parallel generation of hundreds of variables as a distinct capability.
2

Internet Data: A Larger Model Does Not Automatically Acquire Data That Never Appeared Online

A fundamental reason today’s large models are so strong at text, images, video, audio, and code is that these data are abundant on the internet. But much of the truly valuable data in human economic activity has never appeared publicly on the internet: millisecond-level factory sensor curves, equipment control signals, yield and process parameters, internal orders and inventory changes, real trading constraints, fault precursors, failed experiments, and unwritten expert operating knowledge all remain inside corporate databases, machine controllers, laboratories, and tacit human knowledge.

Internet-scale pretraining therefore does not cover “all human experience,” but only the part of human experience that people are willing, able, and suited to express publicly in digital form. This is especially clear in manufacturing: much of the highest-value data has never been uploaded and, because of trade secrets, real-time requirements, device protocols, and sampling costs, will not naturally become a public corpus that can be crawled from the web.

The extraordinary math and coding ability of current models illustrates the same point from the other direction: these capabilities did not simply emerge from crawling more internet data. Model developers have invested heavily in human demonstrations, preference labels, expert design, cold-start examples, verifiable rewards, synthetic data, and automated filtering to actively produce training signals. OpenAI has disclosed that InstructGPT alignment fine-tuning used about 20,000 hours of human feedback; DeepSeek-R1 explicitly used cold-start data, multi-stage training, and verifiable reinforcement-learning rewards to strengthen math, coding, and reasoning.

A core dimension of future AI competition may therefore shift from “who can crawl more internet data” to “who can enter real operations and continuously generate high-value training data.” Industrial equipment, market transactions, robot actions, and scientific experiments may become some of the most important data sources for the next generation of AI.
3

Representing the World: Tokens Represent “What a Symbol Is” Better Than “What a Numerical Change Means”

Language models do not directly understand “20.0,” “20.1,” and “21.0” as three positions on a continuous number line. Numbers are first split into discrete tokens like words and then mapped into embedding space. This works well for language, but it can break some of the most important structure in real numerical spaces:magnitude, distance, continuity, derivatives, periodicity, and the physical meaning of small changes. Research has repeatedly found that simply changing how numbers are tokenized and represented can significantly alter arithmetic and numerical reasoning performance.

This is not a minor issue in production and manufacturing. Equipment control is usually not about what the word “temperature” means, but how pressure, current, efficiency, and the next control action change when temperature moves from 24.80°C to 24.86°C. Market forecasting is not about the concept of “sales,” but the coupled relationships among small changes in sales, price, inventory, and promotions over continuous time.Compressing continuous states into language-style tokens turns smooth numerical neighborhoods into discrete symbolic relationships.

At a deeper level, internet models mainly learn “how people describe the world.” Real production problems require models to learn how the world itself changes: if action a is taken, how does state s move to the next state s'? If ten control variables change together, how does the objective change? Which small changes move the system toward stability, and which make it unstable? This is not a capability that automatically emerges from more text data or a longer context window.

So the third boundary concerns both representation and the relationships among states and variables: A model needs representations of continuous values that respect physical scale, and it also needs to learn the dynamic relationships among actions, state changes, and outcomes before it can truly enter manufacturing, energy, robotics, and scientific control.
Internet AI: text / images / video / audio → Real problems: state / action / outcome / objective → Broader AI: control / robotics / markets / science

From “Generating the Next Token” to “Modeling the Problem Directly”

If a problem is not naturally sequential, there is no need to force it into tokens first. This leads to another path we are exploring:Latent Problem Modeling (LPM). It is neither a Transformer nor token-by-token autoregressive generation. Instead, it tries to learn the latent structure among objectives, solutions, and inputs directly.

Autoregressive Transformers vs. Latent Problem Modeling Transformer / Sequence Generation Latent Problem Modeling Generate output step by step Generate hundreds of tokens or variablesin parallel at once T₁ T₂ T₃ Latent variable z latent space Objective v Solution w Input x
The point of Latent Problem Modeling is not to rename a network architecture, but to change the unit of solving from sequential tokens to the relationships among variables, solutions, and objectives inside the problem.

Latent Problem Modeling Is Well Suited to Four Classes of Problems That Remain Difficult Today

1) Text and Image Generation: From Step-by-Step to One-Shot Generation

Our current result: For text and image generation, Latent Problem Modeling can be about 300× faster than current Transformer autoregressive or diffusion generation methods. The core reason is that hundreds of tokens or variables can be generated in parallel at once instead of token by token or denoising step by step.

Why it matters: Sequential dependence is a bottleneck for real-time LLM deployment. Even acceleration methods such as speculative decoding report mainly about 2–2.5× speedups on a 70B Chinchilla model. On the image side, OpenAI notes that diffusion models typically require tens to hundreds of sequential sampling steps; continuous consistency models achieve about a 50× wall-clock speedup with two-step sampling. If roughly 300× parallel generation can generalize, the impact would extend beyond model architecture to GPU utilization, inference-server design, data-center planning, and the industry logic of buying throughput by adding more compute.

2) Air-Conditioning Efficiency: Taking AI from Internet Modalities into a Continuous Numerical World

Our current result: In air-conditioning control, we achieved about 30% energy savings, overcoming a limitation of AI systems built mainly around internet modalities such as text, images, video, and audio: weak sensitivity to continuous numerical problems. This result has been achieved in a real experimental environment.

Why it matters: Air conditioning is not a small application. An IEA analysis in 2026 estimated global space-cooling electricity use at about 2,900 TWh, with cooling accounting for 14% of global electricity-demand growth since 2015. Even a few percentage points of improvement in control algorithms can carry enormous energy and economic value.

3) Fusion Control: Turning Steady-State Duration into a Continuously Optimizable Problem

The problem we want to solve: Use Latent Problem Modeling to learn the relationships among states, control variables, and steady-state objectives in fusion control, continuously optimizing how long plasma can remain stable. The goal is not a static answer, but whether a fast, dynamic physical system can remain in its target state longer and more reliably.

DeepMind and EPFL have already used deep reinforcement learning to control 19 magnetic coils on the real TCV tokamak at 10 kHz; in 2025, EAST achieved a record 1,066 seconds of steady-state high-confinement plasma operation. At the same time, AI itself is becoming a major consumer of electricity. This could create a positive spiral:Stronger AI improves fusion control → more stable and usable clean energy → energy supports larger-scale AI.

4) All-Modal Robotics: Putting Sensors, Actions, and External Knowledge into One Model

Our direction: Put robot vision and other sensor states, action sequences, task objectives, and external text and knowledge into a single latent problem model. The goal is not for robots merely to replay trained actions, but to build a unified model that can generalize across settings and perform everyday tasks.

Frontier work is already moving toward this goal, but modalities are still joined in different ways. RT-2 jointly trains on internet visual-language knowledge and robot trajectories, encoding actions as text tokens; V-JEPA 2 first learns the physical world from large-scale internet video and then uses a small amount of robot data to form a world model for zero-shot planning and control. The next key step is to move from “joining” sensors, actions, and knowledge to jointly modeling them in a unified problem space.

The next AI curve may not come from a larger version of the same model. It may come from re-modeling the structure of the problem itself: moving inference from sequential to parallel, extending data from internet modalities to continuous values, expanding generation from content to control, and moving robotics from single-mode perception to all-modal world models.

One Important Addition: Two AI Growth Paths Will Develop at the Same Time

These are not product categories called “agents” and “world models,” but two more fundamental directions of growth:The first makes intelligence longer-horizon and more automated within internet modalities. The second moves beyond internet modalities into a real world made of continuous values, sensors, actions, and physical feedback.

Path 1: Long-Horizon Intelligence Within Internet Modalities—Larger Models, Agent Economies, and RSI

White-collar work involving text, code, web pages, documents, email, spreadsheets, and software tools can be organized as long-context processes of information handling, judgment, and execution. As models and contexts grow, what AI first gains is long-horizon capability within internet modalities: a single run can cover more material, more complex chains of tasks, and more complete business processes.

Our view is that within the next 2–3 years, AI’s impact on white-collar work will move from assisting individual tasks to systematic replacement of roles and workflows. When organizations assess AI capability, the unit will no longer be only “what can one prompt do?” It will increasingly map onto conventional organizational concepts such as positions, job responsibilities, workflows, and interfaces between functions.

Single-turn Prompt → Long-horizon Agent Task → Role Responsibilities → Cross-role Workflow → Team / Function

Clear signals are already visible: Claude Sonnet 4.6 supports contexts up to one million tokens; OpenAI’s 2026 work research shows Codex shifting from short interactions to delegated tasks lasting minutes or hours, including in nontechnical functions such as legal and recruiting; METR’s long-term evaluations show continued growth in the “human-equivalent task duration” frontier agents can complete.

Agent economies and recursive self-improvement (RSI) are not a separate path; they are extensions of this same path. Agents connect long-context capabilities to browsers, code execution, enterprise software, and digital workflows, allowing models to act continuously in the internet world. RSI goes further by involving AI in coding, debugging, model training, and AI research itself, creating a self-reinforcing loop of digital capability.

But the boundary of this path is equally clear:Today’s agent economy, and RSI built on coding ability, model training, and AI research, still remain within internet modalities. They can make tasks involving text, code, web pages, and software environments longer and more autonomous, but longer chains do not automatically provide understanding of industrial continuous variables, real sensors, or the consequences of physical actions.

Anthropic’s published work on recursive self-improvement also focuses on coding, debugging, and AI research experiments; METR explicitly cautions that its long-horizon agent evaluations mainly come from software engineering, machine learning, and cybersecurity and cannot be directly extrapolated to all real-world work. These advances are important, but they show more precisely that “digital intelligence” within internet modalities will continue to grow rapidly—not that it has already crossed the boundary into the physical world.

Path 2: Move Beyond Internet Modalities—Gain the Ability to Solve Problems in the Real World

The second path is not to make the same digital intelligence longer, but to move AI into fundamentally different data structures and relationships among variables. In manufacturing, energy, robotics, transportation, and scientific experiments, the core objects are often not sentences and images but temperature, pressure, current, speed, inventory, actions, positions, material states, and experimental parameters—continuous variables that require the model to face a continuously changing system rather than a static information input.

The key structure of these problems is a closed loop:Observe current state → choose an action → action changes the real world → receive a new state and feedback → decide again. Adjusting a valve changes pressure and temperature; changing compressor frequency changes energy use and comfort; moving a robot arm changes the positions of objects and the environment. AI gains true problem-solving capability beyond internet modalities only when it can learn how actions change the world.

This is why frontier research increasingly distinguishes “digital-world intelligence” from “physical-world intelligence.” When Google DeepMind introduced Gemini Robotics, it explicitly noted that multimodal reasoning over text, images, audio, and video had largely remained in the digital domain; entering the physical world also requires embodied reasoning and the ability to act safely. Meta’s V-JEPA 2 further combines internet-scale video with a small amount of robot action data, enabling a world model to predict how the world changes after an action and use that ability for zero-shot robot planning and control in new environments.

Industry is moving in the same direction. NVIDIA calls this field Physical AI, and treats real-world data, simulation, and world foundation models as core infrastructure for learning in robotics and autonomous driving. Its Cosmos platform is designed to “predict future states, generate physical-world data, and train real-world agents.” Together, these directions show that moving beyond internet modalities requires new training data, new model structures, new validation methods, and closed-loop interaction with real equipment and physical systems.

Latent Problem Modeling is also our exploration in this direction. It does not require every industrial or real-world problem to be translated into natural-language tokens first; instead, it directly models the relationships among states, inputs, candidate solutions, objectives, and continuous variables. Air-conditioning control, fusion steady-state optimization, and all-modal robotics are examples of the problems on this path. At the same time, because of its improved generation efficiency, Latent Problem Modeling can also disrupt the Transformer technology of Path 1 within internet modalities.

Two AI Growth Paths: Longer Within Internet Modalities, Broader Beyond Them Path 1 · Go Longer WithinInternet Modalities Path 2 · Go Broader BeyondInternet Modalities Larger Models · Longer Context Agent Economy · Digital Tool Use RSI · Code / Training / AI-Research Loops White-Collar Role and Workflow Automation Longer and more autonomous,but still mainly digital Sensors · Continuous States · Real Equipment Action → World Changes → New State World Models · Physical AI · LPM Manufacturing · Energy · Robotics · Science Capability enters real-world relationshipsamong states and variables Capability Boundary
Path 1 asks how long AI can operate continuously in the digital world; Path 2 asks whether AI can understand and change the real world.
Section 03 Summary: Path 1 will continue to advance rapidly: larger models, longer contexts, agent economies, and RSI will push white-collar work in internet modalities toward role-, workflow-, and organization-level automation. But they remain on the same “digital intelligence” path. Path 2 truly moves beyond internet modalities and into real-world loops of state, action, and feedback.The first restructures knowledge work; the second opens new problem spaces in manufacturing, energy, robotics, and the physical sciences.
04 · From Problem Space to a Future Industry

If AI Moves from “Generating Content” to “Solving Problems,” We Need to Rethink Problem Space

We usually start by asking what AI can solve, but rarely ask how large the space of problems facing humanity really is. The problems humans can solve today are not most problems; they are only a small subset.

Human-Solvable, Human-Recognizable, and Unknown Problem Spaces Problems Humans Have Not Recognizedor Cannot Recognize Problems Humans Can Recognize Problems Humans Can Solve The small region coveredby current human ability Solvable ⊂ Recognizable ⊂ Unknown / Unrecognizable

Problems Humans Can Solve: problems we can define, understand, and solve with relatively reliable methods;Problems Humans Can Recognize: problems we know exist and whose importance and basic structure we can understand, but for which we lack reliable solution methods. Beyond these lie problems not yet discovered, described, or defined by humans—or perhaps beyond current human cognition.

The most important conclusion: Problems humans can solve are only a small subset of all problems. The truly vast space lies in problems we recognize but cannot yet solve, and in problems we have not even recognized.

The History of Physics Is a History of Recognizable Problems Becoming Solvable

A new understanding of the world does more than add knowledge. Each expansion of theoretical boundaries moves some previously indescribable, incalculable, or non-engineerable problems into the solvable range—and eventually creates new industries.

Newtonian Physics, Quantum Physics, Relativity, and Industry Expansion Newtonian Physics The macroscopic worldbecomes calculable Machinery · Railways · Automobiles Engineering · Aerospace Quantum Physics The atomic scalebecomes designable Transistors · Semiconductors · Microchips Lasers · Computing · Communications Relativity High speed, gravity, andprecision time become calculable GPS · Satellite Navigation · Logistics Ride-Hailing · Network Synchronization

Newtonian Physics: The Macroscopic World Moved from Experience to Calculation

Expansion of understanding: Unified laws of motion and mechanics made trajectories, forces, mechanical structures, and motion outcomes quantitatively predictable.

Newly solvable problems: Mechanical transmission, machine design, structural analysis, railways and transport, engineering manufacturing, flight, and orbital motion gradually moved from craft experience to calculable engineering.

Business forms created: Large industrial systems in machinery, machine tools, rail transport, construction, automobiles, and aerospace.

Quantum Physics: The Atomic Scale Moved from “Classically Unexplained” to “Designable”

Expansion of understanding: With quantum theory, atomic energy levels, electron behavior, and microscopic interactions between matter and light became precisely calculable.

Newly solvable problems: Controlling electron flow, designing material band structures, generating and manipulating coherent light, and storing and processing information at microscopic scales.

Business forms created: Transistors, semiconductors, microchips, lasers, modern electronics, computers, communications, and optoelectronics.

Relativity: High Speed, Strong Gravity, and Precision Time Moved Beyond Newtonian Approximation

Expansion of understanding: Relativity redefined the relationships among time, space, speed, and gravity, allowing high-speed motion, strong gravity, and high-precision clock synchronization to be described correctly.

Newly solvable problems: Satellite time synchronization, global navigation accuracy, and common time references for long-distance communication and positioning.

Business forms created: GPS and satellite navigation in turn support map navigation, logistics dispatch, ride-hailing, mobile location services, and communications-network synchronization. NIST notes that GPS satellite clocks accumulate about 38 microseconds per day of combined special- and general-relativistic offset, which must be corrected.

New understanding of the world → More problems can be described → More can be calculated → More can be engineered → New industries emerge

This changes how we should think about AI. Its value should not be reduced to “doing what humans already know how to do, only faster.” If AI can turn more recognizable but unsolved problems into calculable ones, then the real commercial upside comes from expanding the set of problems humans can solve.

With AI, the Solvable Problem Space Is No Longer Limited to Human Ability

Tool AI and autonomous AI push the boundary of solvable problems outward in different ways. Tool AI still relies on humans to pose the problem, define objectives, and judge results; autonomous AI begins to discover intermediate problems, form hypotheses, design experiments, search solution spaces, and iterate on objectives on its own.

Humans, Tool AI, Autonomous AI, and Problem Space Problems Humans Have Not Recognizedor Cannot Recognize Problems Autonomous AI Can Solve Problems Humans Can Recognize Problems Tool AI Can Solve Problems Humans Can Solve The smallest region covered byunaided human problem solving Human-solvable ⊂ Tool-AI-solvable ⊂ Human-recognizable ⊂ Autonomous-AI-solvable ⊂ Unknown / Unrecognizable

Human-solvable problems: Tasks humans can reliably complete using their own knowledge, experience, and existing tools.Tool-AI-solvable problems: Humans still define the problem, but AI makes some previously too costly, inefficient, or computationally large problems solvable.Human-recognizable problems: We know they exist, but lack reliable methods to solve them.Autonomous-AI-solvable problems: AI begins to discover intermediate problems, form hypotheses, design experiments, and iterate objectives autonomously. Beyond that still lie problems humans have not recognized or cannot recognize.

AI’s most important long-term value is not replacing humans in more known tasks, but continually expanding the boundary of what problems can be solved.

In 20 Years, AI May Evolve from Selling Models to Trading Problems and Outcomes

Today, the main things traded in AI are models, compute, software, and services. In 20 years, the core object of exchange may no longer be “access to a model,” but the problem itself and the outcome of solving it.

A Future Platform for Trading and Matching Problems People Who HaveProblems Companies · Research InstitutionsGovernments · Individuals Objectives · Context · DataOutcome Value Problem Trading andMatching Platform Standardize Problems · Verify ObjectivesPrice Them Match Algorithms · AI · ExpertsData · Compute Settle Based on Results People and AI That CanSolve Problems Researchers · EngineersDomain Experts Startup Teams · Specialized AIAutonomous Agents

At that point,the object of exchange will shift from buying software / APIs / compute to buying solved outcomes;pricing will shift from seats, tokens, and GPU-hours to problem difficulty, outcome value, verification cost, success rate, and even value sharing;organizational boundaries will also change—companies will not need to permanently employ teams for every possible problem. They could publish problems to a global problem market, where the most suitable people, specialized AI systems, or autonomous agents compete to solve them.

AI roles will diverge further: tool AI improves the efficiency of existing problem solvers; autonomous AI may become an independent solver in the market, continually discovering problems, proposing solutions, verifying results, and building a reputation.

Software Trade → Model/API Trade → Agent Services → Problem-Solving Capability Trade → Problem and Outcome Market

Why Could This Industry Be So Large? A “Tax-Base” Style Estimate

Here, “tax” is not a legal tax, but a platform-economy structure:A platform occupies a sufficiently fundamental, frequent, and hard-to-bypass exchange relationship between supply and demand, then captures a small but stable share of each exchange. The simplest example is not search or social media, but trade itself: a platform connects people who want to sell with people who want to buy, provides search, trust, payment, fulfillment, and traffic, and continuously captures part of the value from transactions and merchant activity. E-commerce platforms such as Taobao and Amazon are the clearest prototype of this transaction-platform model.

Industry revenue potential ≈ Platform-addressable “tax base” × Transaction penetration × Platform take rate
From Commerce Take Rates to a Problem-Solving “Tax” Commerce PlatformsTax base: product supply, demand,and transactionsTaobao / AmazonCommissions / MerchantServices / Ads Search EnginesTax base: information demandGoogle / BaiduAdvertising Social NetworksTax base: social relationshipsFacebook / WeiboAdvertising Short VideoTax base: spare time / attentionTikTok / DouyinAds / E-commerce AI ToolsTax base: cognitive work / decisionsChatGPT / DoubaoSubscriptions / Ads/ Commissions Problem MarketTax base: value of alltradable problemsHuman Experts + Companies+ Autonomous AISuccess Fees /Value Sharing
Platforms begin with simple product transactions and gradually occupy more abstract, more frequent, and broader forms of human exchange.

Commerce Platforms: Trade Itself Is the Simplest Form of Problem and Supply-Demand Matching

Taobao and Amazon connect one of the oldest and most intuitive supply-demand relationships: one side has goods and the ability to sell, while the other has purchasing demand. The platform does not need to produce every item. Through search and recommendations, merchant access, reputation systems, payment, warehousing, logistics, and after-sales service, it lets dispersed buyers and sellers transact at low cost.

Its “tax base” is the merchandise transactions and merchant activity on the platform. In 2025, Amazon’s third-party seller services revenue was about $172.16 billion. The company states that seller programs can charge fixed fees, percentages of sales, or per-item fees. In Alibaba’s FY2025, Taobao and Tmall Group customer management revenue grew 6% year over year, supported by growth in online GMV and take rate. This is the classic platform structure:you do not need to own all the goods; you only need to become the underlying transaction layer between product supply and consumer demand.

From this perspective, the business structure of later search, social, short-video, and AI platforms did not fundamentally change. What changed was the object that could be platformized and monetized: from goods to information, social relationships, attention, cognitive work, and eventually perhaps the outcomes of solving problems.

Search Engines: “Taxing” Information Exchange

Google and Baidu connect people who provide web pages, products, and knowledge with people looking for information. Search engines do not produce most internet information, but they control the gateway through which people find it. Their tax base is global information demand, monetized mainly through advertising by click, impression, or conversion. Google’s advertising revenue was about $294.7 billion in 2025; Baidu still generated substantial online-marketing revenue in 2025.

Social Networks: “Taxing” Human Social Relationships

Weibo and Facebook / Instagram digitize and scale the human needs for expression, display, relationships, and group belonging, then insert advertising into the exchange of social relationships and information flows. Meta’s total revenue was about $201 billion in 2025, with the company reporting that growth was mainly driven by advertising; Weibo generated about $1.5 billion from advertising and marketing in 2025.

Short-Video Networks: “Taxing” Spare Time and Attention

Douyin and TikTok match users’ abundant fragments of time with creators’ ability to capture attention. Recommendation algorithms connect the two with unprecedented efficiency, then monetize through advertising and e-commerce. Reuters reported in 2025 that TikTok’s U.S. business was expected to generate as much as about $25 billion in annual revenue.

AI Tools: Beginning to “Tax” Cognitive Work and Decisions

AI tools such as ChatGPT and Doubao connect people’s cognitive needs at work and in daily life with models’ ability to replace part of reading, searching, writing, analysis, planning, and decision-making. This goes beyond search: users are no longer just finding information; they are handing part of the thinking process to AI. Monetization is also moving from simple software fees toward subscriptions + advertising + transaction commissions.

Problem Markets: Ultimately “Taxing” Problems Humanity Needs Solved

Further ahead, the largest tax base may not be information, social interaction, attention, or even cognitive work, but all the problems in human economic and social life that are worth solving. Those who have problems provide objectives, data, and verifiable outcomes; people, companies, or AI systems capable of solving them provide solutions. The platform handles problem specification, capability matching, verification, trust, and settlement, then captures a share of the real value created by solving the problem.

So a growth estimate for 20 years from now should not focus only on how many tokens or subscribers exist today. It should ask:How much value does society create each year because problems are solved better? If a platform can become infrastructure for the exchange of problems and solutions—just as Taobao and Amazon carry commerce, Google carries information exchange, and social networks carry relationship exchange—even a small take rate could create an industry larger than today’s search, social, short-video, and AI-tool businesses.

Conclusion · The AI Industry 20 Years From Now

For the past few years, the AI narrative has often started with: “I have a large model—what else can it do?” The more important question for the next phase may be the reverse:I have an important problem—which model is best suited to solve it?

What is truly scarce may not be the chat box itself, but unique data, real feedback, domain understanding, objective definition, and the ability to enter a verifiable problem-solving loop. Industrial control, robotics, energy, markets, and scientific experiments are all waiting for problem-modeling approaches better suited to them.

Models Will Get Cheaper, Problems Will Get More Valuable

Twenty years from now, the AI industry may no longer center on who owns the largest model, but on who can solve important problems faster, cheaper, and more reliably. Models may become infrastructure, agents may become labor, and the largest new industry may emerge from a transaction network connecting problems, solutions, and outcomes—a new industry centered on problem-solving capability.

Comments

Popular posts from this blog

A Perplexity Benchmark of llama.cpp

Serving Llama-2 7B using llama.cpp with NVIDIA CUDA on Ubuntu 22.04

Thoughts on AIGC for Non-AI Industries