The Hard Problems in Decentralized AI Nobody Has Solved Yet
TL;DR: Decentralized AI faces four unsolved problems: governing a network without it being captured, evaluating output quality without a central referee, measuring contribution in a way that cannot be gamed, and matching centralized latency and reliability. None has a general solution today. The useful signal when evaluating any project is whether it names these honestly or talks around them.
Key Takeaways
- Governance is hard because the mechanism that distributes power can also concentrate it in whoever holds the most tokens.
- Evaluating output quality without a central referee is unsolved in the general case, since benchmarks get gamed once they carry economic weight.
- Measuring contribution rewards whatever is easiest to measure, which is rarely the same as what is most valuable.
- Latency and reliability decide adoption, because users will not accept a worse product for a better ownership structure.
- The signal of a serious project is whether it names unsolved problems or talks around them.
Most writing about decentralized AI is either promotional or dismissive. The promotional version treats every hard problem as already handled. The dismissive version treats those problems as proof the whole idea is unserious. Both are less useful than simply naming what is actually difficult. Perspective AI is built on the view that decentralized AI is worth doing precisely because the problems are real, and that a project earns trust by being specific about which ones it has not solved.
Here are the four that matter most.
Governing a network without it being captured
The obvious way to govern a decentralized network is token weighted voting: hold more tokens, get more say. The problem is that this reproduces the exact concentration decentralization was supposed to prevent. Whoever accumulates the most tokens governs, and capital is easier to accumulate than merit.
The alternatives have their own failure modes. Reputation systems reward demonstrated contribution rather than holdings, but they are slow to bootstrap and tend to entrench whoever showed up first. One person one vote sounds fair until you try to prove that participants are distinct people, which is its own unsolved problem. Delegated models concentrate power in delegates, which is representative government with all of the capture risk that implies.
There is no settled answer here. What separates approaches is which failure they choose to accept and whether they admit they are choosing. Timelocks, veto rights, and caps on voting weight all reduce capture risk without eliminating it, and each adds friction somewhere else.
Evaluating output without a referee
Centralized providers decide what counts as a good answer. They run internal benchmarks, employ human raters, and set quality bars. That is a real function, and removing the central authority does not remove the need for it.
The difficulty is that evaluation becomes an economic target. A benchmark with no money attached measures capability reasonably well. The moment rewards depend on that benchmark, it stops measuring capability and starts measuring the ability to score well on that benchmark. This is not hypothetical. It is what happens reliably whenever a metric becomes a payout condition.
Peer evaluation is the usual proposed fix, but it inherits the incentives of the peers doing the evaluating. If evaluators are paid, they have reason to collude. If they are not paid, participation skews toward whoever has a stake in the outcome.
What does work today tends to be narrow. Verifying that a specific computation was performed correctly is tractable, because correctness is checkable. Verifying that an answer was good is not, because quality is a judgment. Projects that conflate the two are claiming more than they have.
Measuring contribution without rewarding the wrong thing
If rewards should follow contribution, contribution has to be measured, and measurement systems reward whatever is easiest to measure rather than whatever is most valuable.
Serving inference is the clearest case, because the work is specific and checkable. An operator either served the request or did not. This is why work performed for the network is a more defensible basis for rewards than participation in general, and it is the reasoning behind treating POV as a work token rather than a utility token. The tradeoffs across different designs are the subject of how token models create or destroy incentives.
Other contributions are harder. Data is valuable but its value is only apparent after training, which is long after you need to price it. Model improvements are valuable but attribution is murky when many changes interact. Governance participation is valuable but trivially farmed by anyone willing to vote on everything.
The honest position is that some contributions can be measured well today and others cannot, and a system claiming to price all of them fairly is overstating what is possible.
Matching centralized latency and reliability
This is the problem that decides whether any of the others matter, and it gets the least attention because it is unglamorous.
Inference is interactive. Users notice hundreds of milliseconds. A centralized provider runs inference in its own data centers on hardware it controls, with no coordination overhead. A distributed network has to route the request, select an operator, verify the work, and settle payment, and each step costs time. The same physics constrains decentralized GPU networks competing with the big cloud providers, and it becomes more acute for distributed training of frontier models, where bandwidth between nodes dominates.
Reliability compounds it. One provider with redundant infrastructure has a failure model it can reason about. A network of independent operators has to assume some will be slow, offline, or dishonest at any moment, and deliver a consistent experience anyway.
This is why serious projects run hybrid infrastructure while the network matures rather than shipping a worse product on principle. People will not accept degraded quality in exchange for better ownership, and asking them to is how good ideas fail in practice.
How to judge progress
None of these has a general solution. That is not a reason for pessimism, it is a description of where the work is. What it does give you is a way to evaluate any project in this space:
- Does it distinguish between what is live today and what is planned?
- Does it name which problems it has not solved, or imply all of them are handled?
- When it describes a mechanism, does it say what the mechanism costs?
- Does it explain what would have to be true for the next stage to work?
A project that answers those clearly is doing real engineering. A project that answers them with confidence about the future is selling a story.
That is the standard Perspective Labs holds itself to. The product is live, the inference behind it is hybrid while the network matures, node provided inference is the destination rather than the current state, and the problems above are open. Saying so is not a weakness in the pitch. It is the only version of the pitch that can be checked, and being checkable is the entire point of building AI that its users own.
FAQ
What are the biggest unsolved problems in decentralized AI?
Four stand out. Governing the network without it being captured by whoever holds the most tokens. Evaluating whether model output is good without a central referee to decide. Measuring contribution in a way that rewards real work rather than whatever is easiest to fake. And matching the latency and reliability of centralized providers, which is the problem that decides whether anyone actually uses the thing.
Why is governance harder in a decentralized AI network?
Because the mechanism that distributes power can also concentrate it. Token weighted voting gives control to whoever accumulates the most tokens, which reproduces the concentration decentralization was meant to prevent. Reputation systems avoid that but are slow to bootstrap and can entrench early participants. There is no settled answer, only tradeoffs.
How do you evaluate AI output quality without a central authority?
This is genuinely unsolved in the general case. Benchmarks can be gamed once they carry economic weight, and peer evaluation inherits the biases and incentives of the peers doing the evaluating. Approaches that work today tend to be narrow, verifying that a specific computation was performed correctly rather than judging whether an answer was good.
Why is latency a serious problem for decentralized AI?
Inference is interactive. People notice hundreds of milliseconds, and they will not accept a worse product in exchange for a better ownership structure. A distributed network adds routing and coordination overhead that a single provider's data center does not have, so closing that gap is a precondition for adoption rather than a detail to optimize later.
How can you tell a serious decentralized AI project from marketing?
Look at how it handles the hard parts. A serious project names which problems remain unsolved, says which components are live today versus planned, and explains what would have to be true for the next stage to work. A project that presents its end state as already shipped is asking for trust it has not earned.
Own your AI instead of renting it
Perspective AI gives you private, user owned AI on a decentralized network, with the parts that are still ahead named honestly rather than implied.
Launch App →