KGeN: The Verified Human Data Network for the AI Era

1. AI Is No Longer Sated by Data
For the past few years, the formula driving progress in AI was surprisingly simple: pour ever larger models, ever more data, and ever more compute into the problem. As models and datasets grew, performance improved almost without exception, and over little more than five years the halting early models evolved into today's systems that hold a conversation. So the entire industry ran down the same road. Gather more, scale more. For a while, that road looked like it had no end.
But that formula has begun to hit a wall, because the rate at which AI consumes data is catching up to the total volume of data humanity has accumulated across the internet.

Source: EPOCH AI - Will we run out of data? Limits of LLM scaling based on human-generated data
The research institute Epoch AI estimates the stock of high-quality, human-generated text on the internet at roughly 300 trillion tokens, and projects that if current training trends hold, this resource will be effectively exhausted sometime between 2026 and 2032, with a median estimate of 2028. Compute keeps expanding on the back of advances in hardware and algorithms, yet the data to feed it has run up against a limit. It has become clear that data is not an infinite resource but a finite one that shrinks the more it is used.
This ceiling has quietly but fundamentally shifted the direction of the AI race. The question is no longer how much data can be gathered, but which data. The Second International AI Safety Report, published in February 2026, puts numbers to the problem. Because most training data originates in the West, model performance skews heavily toward English and Western culture: in one evaluation, an AI model answered 79% of questions about everyday American culture correctly but only 12% of questions about Ethiopian culture. Another evaluation, spanning 83 languages, found performance dropping sharply for languages that use non-Latin scripts or have little digital material. What is scarce is not the sheer volume of data, but suitable data that reflects the true breadth of the world.
So why not have AI simply generate the missing data itself? Synthetic data is indeed cited as a leading candidate, but it carries a well-known trap. When a model is trained on data that models themselves produced, each successive generation shows the same pattern: the diversity of outputs collapses and biases amplify. This is what is known as model collapse. Train AI on nothing but its own output, over and over, and performance eventually degrades. One conclusion follows. The more bots and synthetic data fill the internet, the harder it becomes to tell the real from the fake, and the scarcer data from verifiable real people becomes. It becomes a resource whose value rises with its scarcity rather than one whose value falls with abundance.
Moreover, the next candidates to fill the void left by exhausted text, namely image, video, and motion data, demand human hands even more deeply. Collecting whatever is scattered across the web is not enough for this kind of data; it becomes useful only when real people record and refine it themselves in well-designed settings. What AI needs to reach its next stage, then, is not more data but data produced by diverse, real, verified humans.
Yet that kind of data is the hardest to gather. Collecting the traces that people around the world leave behind in their own distinct environments, and doing so in a form that confirms each person is genuine, is an entirely different problem from simply increasing volume. It requires three things at once: the scale of people spread as wide as the world itself, the verification that confirms they are real people, and the diversity that does not tilt toward any one side. Taken separately, each is only half an answer. A vast but unverified user pool cannot guarantee trust, because there is no way to know how many bots and fakes are mixed in; a small, rigorously vetted group is too narrow to hold the diversity of the world. A human network that satisfies scale, verification, and diversity all at once: here lie the conditions for the data supply AI needs to reach its next stage.
2. KGeN, a Verified Human Network
KGeN is a project that positions itself as a human network embodying all three of these conditions at once: sufficient scale, verifiability, and diversity that captures the breadth of the world. KGeN describes itself as Verified Human Infrastructure, and it has gathered 61.9M users across more than 60 countries, along with a system for verifying that each of them is a real, unique individual.
2-1. How KGeN Verifies Humans: The Verified Reputation Engine
KGeN began as a verified distribution layer, infrastructure built to solve a problem that DeFi and consumer applications share: identifying genuine users and actually reaching them. Any service trying to acquire new users runs into the same wall. Offer rewards and people flood in, but a large share of them are duplicate accounts or bots that collect the reward and vanish. Sorting out who is a real person, who is a duplicate, and who is a bot has been KGeN's task from the very start. KGeN calls this system for verifying the authenticity and activity of each individual user its Verified Distribution Protocol, or VeriFi.
At the heart of VeriFi is a reputation-scoring system called the Verified Reputation Engine. It measures a user across five parameters. Proof of Human confirms the authentic human presence behind the account, the foundational layer of trust. Proof of Participation gauges the extent and intensity of a user's interaction with quests, campaigns, and the KGeN platform. Proof of Skill evaluates their proficiency across tasks and data domains. Proof of Commerce captures their economic activity and transactions within the KGeN ecosystem. Proof of Social Network maps their social ties and influence within the community. Each parameter draws on more than 100 sets of attributes per user, and across the entire network the attributes accumulated this way exceed 2 billion.

A user's final reputation is calculated as a weighted sum: each of these five parameter scores is multiplied by a different weight and then added together. The point of a weighted sum is to place greater weight on the parameters that matter most for verification. A bot account that has artificially inflated its Proof of Participation or Proof of Commerce figures still cannot raise its overall score if its Proof of Human sits at the bottom, because of the heavy weight attached to it. This computation happens off-chain, after which the score and its underlying data are recorded together on-chain.
This reputation score lives on-chain as the metadata of a dynamic NFT that holds the user's identity. As activity accumulates and the score updates, the metadata updates with it. Users own this reputation directly and can carry it to other services. KGeN handles the underlying data from the position of a custodian, storing separately the personal information that users provide or that is inferred, records of platform activity, and derived data generated through analysis. What to expose externally, what to make public and what to keep off-chain, and whether to delete data are all decisions left to the user.
The real value of verification lies less in the volume of data than in the fact it proves. Where tools like CAPTCHA or Cloudflare's bot filtering confirm just once whether the current session belongs to a human and then stop, the Verified Reputation Engine proves that the same person has built a consistent record over a long stretch of time. A single ID card and a lifetime of activity are not the same thing. And the better AI becomes at imitating people, the more the value of a continuity it cannot imitate actually rises. This continuity becomes the core differentiator of KGeN's data, as later sections will show.
2-2. The Oracle Network: How to Make the Scores Trustworthy
The reputation score produced through the Verified Reputation Engine is computed off-chain, and if that value stayed inside KGeN alone, external customers would have to trust a score KGeN assigned. In that arrangement, the provider of verification backs the basis for that verification with nothing more than its own claims, the most fragile point in any business built on selling trust. The oracle network resolves this. An independent set of validators recomputes the score KGeN produced, reaches consensus, and records the result on-chain, so that trust in the score is separated from KGeN as a party. This separation is precisely why external customers pay for the data.
The oracle network consists of the validators, called oracles, and Keyholders. Holders of the $KGEN token can acquire an NFT called a Genesis Key, and a Keyholder who holds more than 200 of these keys earns the right to run an oracle node. Tying validation rights to tokens and keys binds validators economically to the honesty of the network.

Each oracle independently computes the reputation score from a user's activity data, then signs and submits it. Once the values submitted by at least 67% of oracles converge on a single result, that consensus value is recorded on the blockchain. Divergent values and repeatedly deviating values are excluded, and an oracle found to be acting in bad faith has its staked assets slashed. Because the staked assets slashed exceed whatever could be gained through manipulation, the incentive to attempt it disappears. The network currently runs on a Proof of Authority (PoA) model, in which KGeN selects the operators; in roughly three years it is set to transition to a Proof of Stake (PoS) model in which anyone can stake tokens to participate. Leading Korean validators and crypto companies, including Cosmostation, b-harvest, KODA, and MarbleX, have joined as KGeN oracle nodes.
Through this structure, the reputation score moves beyond an internal KGeN metric to become infrastructure that outside parties can verify and use. AI companies can use it to assess the authenticity and qualifications of data workers, and Layer 1s can use it as a basis for identity and governance and as a gauge of whether a user is genuine. For users, too, this reputation is an asset that is not bound to any single platform, so they can prove themselves across multiple services with the score they own.
2-3. The Clan Network: Scaling Through Reach and Diversity
The value of a network grows in proportion to the square of the number of participants. Double the users and the number of connections that can flow among them swells nearly fourfold. Whether a telephone network or a social platform, this is why the value each individual enjoys jumps as more people use it. This is Metcalfe's law. Early participants are hard to gather, but once a network crosses a certain size it enters a zone where value amplifies on its own, and from there it grows explosively. KGeN actively uses this law as a go-to-market strategy. Instead of acquiring users one at a time through advertising, it brings already-formed communities into the network wholesale. In this structure, which KGeN calls the Global Clan Protocol, a single community joins as a unit called a clan.

A Clan Chief who leads a community, whether an offline one such as a university club or a workplace group, or an online one on Discord, Telegram, and the like, brings that community in as a single clan. Each clan operates like a DAO, with the Clan Chief recruiting members, organizing activity, and distributing rewards. There is no need to persuade people one by one; hundreds or thousands who are already bound together are onboarded at once, so the network swells exponentially. Clans gathered this way now number more than 30,000 across over 60 countries, and together they form a user base of 61.9M.
This approach produces diversity alongside scale. Because each clan is a group rooted in a real region and culture, a community in Brazil, a university in Korea, a gathering in the Middle East, the network naturally captures a diversity that a uniform acquisition channel could never hold. A structure that brings in local communities as they are captures what replicating users from one country into another would miss. This structure, which produces scale and diversity at the same time, is what sets KGeN's network apart from a mere user pool.
3. Verified Humans Become AI's Data
The place where the network of verified people described in the previous chapter turns into value is the market for AI data. The data exhaustion and the shift toward quality laid out in Chapter 1 are growing this market rapidly.
Among the 61.9M people verified through the Verified Reputation Engine, those whose qualifications and skills in specific fields have been confirmed form an AI expert network of roughly one million people. They are drawn from the network's 60-plus countries, spanning over 30 languages and more than 10 domains. It is a group filtered by reputation from a broad verified base, keeping only those whose qualifications match the task. This business operates as a two-sided marketplace connecting AI companies that need data with verified experts. When a company commissions data, KGeN draws on reputation to select users whose qualifications and track record fit the job, assigns them the task, processes the output into data the company can use, delivers it, and collects a fee.
Consider a hypothetical case: suppose a company is building a model specialized for the medical field. Such a company needs data produced by real, capable people, yet an anonymous crowdsourcing platform cannot guarantee that its workers are even human. This is where KGeN's structure could be applied. Verified users are classified into cohorts by their reputation scores, and when a company submits a task, it is routed to the most suitable cohort. Because these contributors are verified through on-chain reputation, the company knows the work is done by real, verified people rather than anonymous accounts. This kind of data work is done by people reviewing and refining a model's responses, and what separates it from conventional mechanical crowdsourcing that hands work to anonymous users is that the contributors are verified real people.
Now that text models have swallowed nearly all of the internet's language, the next frontier is the physical world. For robots to pick up objects, for devices to read a situation, and for agents to carry out real-world tasks on our behalf, models have to learn how people actually move and act. Yet this kind of behavioral data is nowhere near sufficiently accumulated anywhere on the internet. Egocentric recordings of people cooking, cleaning, and handling tools are obtained only when someone deliberately creates them, and creates them across varied environments at that. This is why the bottleneck for physical AI is said to lie in data rather than in compute or algorithms.
This is exactly the area KGeN invests in most heavily. KGeN defines its AI data as multimodal, spanning text and speech, image and video, and reaching further into motion, and it places the center of gravity on robots and AI that operate in the physical world. That direction shows plainly in the business roadmap. In 2026, KGeN is running a workflow to collect, verify, annotate, and quality-control 22,000 hours of single- and multi-speaker speech data across 33 languages, and then plans to move on to egocentric video for physical AI, standing up data pipelines in residential settings in Latin America and manufacturing floors in South Asia.
In this business, KGeN collects the data, and its exclusive partner Humyn Labs processes that data and supplies it to frontier AI labs. KGeN provides the Human Infrastructure while Humyn Labs provides the Human Intelligence. Working with Humyn Labs, KGeN has in fact already collected more than 20,000 hours of egocentric video in just three months. Capturing the everyday motions of people cooking, cleaning, and handling objects, this video is used to teach robots human behavior.

Source : X(@KGeN_IO)
This is where the network's diversity turns into value. Give the same task to users in India, Brazil, and Saudi Arabia and each performs it differently, because the way people interact, the dialects and expressions they use, their hand movements and gestures, their habits in handling devices, and the whole cultural context surrounding all of it differ from region to region. This behavioral diversity, which existing datasets skewed toward the English-speaking world and developed countries structurally cannot reach, is the most precious raw material for a model that must learn the physical world. And such data cannot be obtained by gathering what is scattered across the web, nor by staging it in a studio. It is obtained only when real people record it themselves in their own daily lives, and this is exactly where KGeN's asset, a network of verified people rooted in their regions and cultures, fits best. Now that public data is running dry and frontier labs have begun pouring tens of millions of dollars into a single specialized multimodal dataset, the position of whoever has secured data produced by verified humans becomes that much clearer. And the real revenue from this data business is the foundation on which KGeN has redesigned its token model.
4. KGeN 2.0: Revenue Becomes Tokenomics
In most token economies, a token's price connects poorly to the project's actual business performance. The business can do well without the token rising, and the token can rise with no relation to the business; both are common outcomes. The recently announced KGeN 2.0 aims to break this disconnect, presenting a structure that uses real revenue from the business to buy back and retire tokens, enhancing value for token holders.
The starting point of KGeN 2.0 is a one-time, large-scale retirement called the Genesis Burn. 22M $KGEN was retired, roughly 10% of the circulating supply. It also set in motion a mechanism that links revenue to buybacks and supply reduction, running in four steps: revenue is generated in the AI data business; a portion of the contribution margin, the profit left after variable costs are subtracted from that revenue, is allocated to fund the reduction; this funding is used to buy $KGEN on the open market; and the purchased tokens are retired to enhance value for holders. It is geared so that as AI revenue grows, the funding for this reduction grows with it.
KGeN runs several lines of business and reported approximately $85.8 million in annual recurring revenue (ARR) as of March 2026. ARR is a metric that tallies, on an annualized basis, revenue that recurs every year, such as subscriptions. The AI data business is one of those revenue lines, the revenue that comes from verified users carrying out AI companies' training-data tasks.
At the heart of KGeN 2.0 is the token reduction, and its starting point is the contribution margin of the AI business. Because the funds used for the reduction are revenue generated by the AI business rather than investor capital or the foundation's reserves, the returns flowing back to token holders grow as KGeN grows, connecting structurally to the token's value.
For this structure to earn trust, the revenue itself has to be verifiable, because a revenue-linked token model only holds up when outsiders can confirm that revenue. To this end, KGeN plans to have its revenue independently audited and published on-chain, and to make its buybacks and token retirements verifiable on-chain as well.
That said, a few things must hold for this structure to work. AI revenue, the basis of the reduction, still accounts for only around 10% of total ARR, a small share, and the phase in which the varied behavioral data KGeN emphasizes, beyond the data already available on the web, becomes necessary has only just begun. In the end, whether this revenue arises steadily and can grow is the crux, and that trajectory is something to watch.
5. Conclusion: AI Will Shift Back Toward the Human
The center of gravity in the AI race is shifting. Today's bottleneck discussion centers mainly on hardware infrastructure such as memory and power. But this bottleneck is moving from hardware to data. However large compute and power grow, performance hits a wall without new data to make models better. As text is exhausted and synthetic data fills the internet, the value of data produced by verified people will rise.
As the frontier crosses into multimodal and physical AI, this trend grows sharper. Physical AI that acts in the real world, like a robot, has to learn from records of people moving and making judgments in real environments. That kind of data has to be made by real people themselves. The importance of data produced by verified humans grows another notch here.
What KGeN has built is exactly that resource. A network of verified people across more than 60 countries, a verification system that inscribes their authenticity and qualifications on-chain, a marketplace that connects the data they produce to AI, and a structure that uses that revenue to buy back and retire tokens and return value to holders are all linked into one.
The AI data business has only just begun to scale, and whether the pace of that growth will actually complete KGeN's vision remains to be seen. Even so, the direction in which verified humans become the core data supply of the AI era is becoming clear, and as one of the networks that has secured the widest access to that data, the vision KGeN puts forward is worth watching.
Disclaimer
I confirm that I have read and understood the following: The information contained in this article is strictly the opinions of the author(s). This article was authored free from any form of coercion or undue influence. The content represents the author's own views and does not represent the official position or opinions of CrossAngle. This article is intended for informational purposes only and should not be construed as investment advice or solicitation. Unless otherwise specified, all users are solely responsible and liable for their own decisions about investments, investment strategies, or the use of products or services. Investment decisions should be made based on the user’s personal investment objectives, circumstances, and financial situation. Please consult a professional financial advisor for more information and guidance. Past returns or projections do not guarantee future results. This article was written at the request of KGeN. All content in this article was written independently by the author(s), and neither CrossAngle nor KGeN had any editorial control or influence over the content. The author(s) may hold the cryptocurrencies mentioned in this article at the time of writing.
Xangle or its affiliated partners own all copyrights of the written or otherwise produced materials and content provided on the platform. Any illegal reproduction of such content, including, but not limited to, unauthorized editing, copying, reprinting, or redistribution will result in immediate legal actions without prior notice.

![[Xangle RWA Series] Compliance](https://resource.xangle.io/files/content/9EB5169D9605BBF00DEC081C07A10955_1784185676780.webp)


