ABSOLUTELY UNLIMITED.

Unlimited on every plan. More room and priority as you go.

Core

6/ month

32K context

Everything you need to start.

  • 32K context
  • Absolutely unlimited inference
  • Standard queue
  • Streaming API

Core

Everything you need to start

Core is your starting point for unlimited inference, with up to 32K of context per request.

Use it for everyday agents, automations, and new ideas.

Put your agents to work and give them no reason to stop. Every request is unmetered.

One Million

25/ month

1M context

Everything in Pro, plus

  • Yes, still unlimited inference.
  • About 15x Pro context, up to 1M
  • Highest queue priority
  • Room for whole codebases and long-running agents

One Million

Built for your biggest requests

One Million gives you room for exceptionally large requests.

Because full-window jobs require much more GPU memory, they may queue longer than ordinary requests.

For the smoothest latency, use 256K or less when possible and save the full 1M window for non-urgent work.

The context is massive, the inference is unlimited, and your requests receive priority access.

Have fun with it!

UNLIMITED ON EVERY PLAN.

Your work stays yours. We never use your prompts or outputs to train AI models. Prompts may be encrypted and retained for a limited period for security, abuse prevention and legal compliance.

* Fair usage applies. Heavy workloads may wait longer during demand spikes, but there is no monthly usage cap. See our Terms of Service and Privacy Policy for full details.

POWERING FEIHOA

UNLIMITED. NOT UNDERPOWERED.

FEIHOA data centers run Qwen3.8 27B FP8 Uncensored from HauhauCS on every plan, with up to one million tokens of context for the most demanding tasks.

Meet Qwen3.8 27B