Put in your monthly request volume, your own provider prices and the share of work each path would take. It returns an estimate you can take apart, never a quote.
Nothing is sent anywhere. The numbers stay in your browser.
Every assumption in play, and which of them are yours to change. The values it opens with are editable examples rather than recommendations.
Estimates rest entirely on the values you enter. They exclude taxes, currency movement, implementation work and any charge not entered. Confirm commercial terms with FenxLabs, and confirm inference rates with the provider that bills you for them.
What ARC charges to route a request: a monthly fee, the calls it includes and the rate above that allowance. The modeller works in that shape so a scenario can be built before a quote exists.
Model inference, tools and every other provider charge sit outside this figure. Read it as the ARC line on your bill, monthly or annualised, and add the rest separately.
Tiers read by the modeller: Pay As You Go, Starter, Pro, Team, Business, Enterprise. Rates under Enterprise are negotiated per organisation, so anything the modeller shows there is a planning assumption.What it costsInference stays a direct relationship between you and your providers, at their rates. We take no margin on it. The modeller works from the prices you enter, against a baseline you define.
ARC does not guarantee that a given share of requests can move to a lower-cost model while holding the quality that work requires. Validate the routing policy and its outcomes on representative workloads before you plan against the difference.
Priced by the number of human users you licence. One annual fee, no metering. The licence is one line of this bill, and your own environment is the rest.
Every figure in this scenario is a planning assumption you enter, including any licence value the modeller opens with. None of them is a FenxLabs quote. Infrastructure, operations and support belong to your environment, so only your team can price them.
The modeller applies no exchange rate, so keep one currency across a scenario.Explore the Self-hosted LicenceARC does not cut spend by sending every request to a weaker model. It cuts spend by not buying capability the request does not need.
Full output quality
For work where being wrong costs more than the request does. ARC pays for headroom wherever the harder answer is worth it.
Little practical difference on routine work
The practical default for mixed work. Efficient and specialised resources carry what they can hold, and the hard requests still escalate.
Lower, and acceptable for these jobs
For high-volume mechanical work. Output quality drops, and the requests that genuinely need a frontier model still get one.
Modelled against routing the same workload entirely to frontier models, across a mixed estate. Past the balanced setting, saving more costs output quality. On bulk classification or first-pass triage that is the right trade. On a customer-facing summary it is not. You set the position per use case, not once for the whole estate.
Past the balanced setting, saving more costs output quality. On bulk classification or first-pass triage that is the right trade. On a customer-facing summary it is not.
You set the position per use case, not once for the whole estate.
Modelled against routing the same mixed workload entirely to frontier models. Your result depends on your providers, traffic and workload mix.See how a request is handledThirty minutes with an engineer. Your numbers. A straight answer.