MEET THEQUEUE
FLAT-RATE INFERENCE
UNLIMITED TOKENS AT A FIXED PRICE.
- Monthly tokens
- Unmetered
- Monthly request cap
- Unlimited
- Active per account
- 1 generation
- Queue priority
- Plan + demand
THE HARDWARE
BEHIND EVERY REQUEST.
Current FEIHOA production runs on NVIDIA's Blackwell architecture. Each GPU processes multiple requests across accounts, with the priority queue deciding what enters available capacity next.
Current FEIHOA production runs on NVIDIA Blackwell. Each GPU handles requests across accounts while the priority queue decides what runs next.
- PER-GPU PEAK PREFILL*
- 3,1003,100PREFILL TOK/S
- PER-GPU PEAK GENERATION*
- 206206GENERATION TOK/S
SHARED CAPACITY. SMART ORDER.
CoreStandardQUEUE PRIORITY
ProPriorityQUEUE PRIORITY
One MillionHighestQUEUE PRIORITY
Plan priority affects queue order.
GET THE MOST FROM FEIHOA.
Every workload is different. Match request volume and context to current demand for the best throughput. See the Terms for queue limits and fair-use details.
- Avoid request floods.Send work at the pace you can use it. Repeated bursts from the same account are deprioritized when capacity is busy.
- Trim unused context.Send only the history and files the task needs. Shorter context uses less GPU capacity and usually finishes faster.
- Build for retries.FEIHOA is designed for agentic work, not instant chat. Expect variable latency and make retries safe so work is not duplicated.
THE MODEL
MEET QWEN3.8 27B
The uncensored model behind FEIHOA, built for long-running agentic work and up to one million tokens of context.
Meet Qwen3.8 27B