Use Inference Pool
A chat app and API over many models. Send a prompt, the router picks eligible GPU supply, and tokens stream back in real time.
Open the appEvery GPU on the network is one pool of supply. Send a request, the router draws from it, and tokens stream back. A Uniswap v4 hook pays the same pool from $INFP liquidity.


The router draws from the pool by model, availability, and measured speed. The worker executes and tokens stream straight back. Deeper the pool, faster the match.
Every worker runs inferencepool worker benchmark on join and reports its own numbers. The router weights routes by measured tokens per second, not by advertised hardware.
Specs from NVIDIA's RTX Blackwell architecture whitepaper. Inference Pool workers report the same fields, VRAM, cores, bandwidth, for routing.
Usage is counted, earnings are credited, and prompt context is discarded once the response completes.
Compute liquidity and token liquidity feed the same place. Holders earn from fees that were actually collected, never from emissions.
Workers join and advertise the models they serve and how fast they serve them. The router draws from that pool per request. Supply is pooled, not assigned, so a deeper pool means a faster match and better fan-out.
$INFP trades in a Uniswap v4 pool with a custom hook attached. On every swap the hook takes 0.30% and routes it straight into the holder reward pool that inference margin already fills. Capped at 1% in the contract, and renounceable.
A chat app and API over many models. Send a prompt, the router picks eligible GPU supply, and tokens stream back in real time.
Open the appRun the worker client, benchmark your machine, list supported models, and accept routed jobs. Earnings are credited per usage.
Become a providerHolders stake for a pro-rata share of the reward pool. It fills from two directions: 80% of every job's margin, and 0.30% of every $INFP swap through the Uniswap v4 hook.
Read tokenomics