Live paid inference

Eight models.
Pay for tokens,
not seats.

OpenAI-compatible inference for agents, automations, and developers. Choose a simple flat rate through Halo or model-specific pricing through AntSeed. Settle in USDC on Base.

6 verified buyersIndependent on-chain buyers
1,041 requestsPublic indexed request history
37 settlementsMachine-native Base settlement
0 / 50 loadCurrent / maximum concurrency

Signed seller metadata · domain + GitHub verified · public activity observed 2026-08-24. Inspect machine-readable proof →

Two ways to buy

Pick the billing model that fits the workload.

Both routes reach the same eight-model fleet. Halo is simplest for mixed workloads; AntSeed gives price-sensitive agents model-level control.

Per-model rateLowest entry price

AntSeed

0.009 USDC / million input tokens, from

Choose a model-specific input and output rate. The provider publishes signed metadata, capabilities, load, pricing, and OpenAI-compatible protocol support for machine discovery.

  • Signed public provider metadata with seven services
  • Separate input, output, and cached-input prices
  • Capabilities declared for all seven AntSeed services
Published-cost calculator

Compare the rails before funding one.

Enter an expected token mix. The calculator reads the same machine catalog agents use and compares Halo's flat total-token rate with AntSeed's model-specific input, output, and cached-input rates.

Cached input must be part of the input-token total. Change any field to recalculate locally; no prompt or wallet data is sent.

Halo—USDC · flat total-token rate
AntSeed—USDC · model-specific rates

Loading published rates…

The machine catalog remains the source of truth.

Published token charges only. Wallet gas, network-specific fees, retries, and future price changes are excluded. Model positioning is a routing hint, not a benchmark guarantee.

Model fleet

One catalog for reasoning, code, speed, and Chinese.

All eight models returned HTTP 200 through the configured OpenAI-compatible upstream. The catalog advertises at least a 128,000-token context window; this is separate from a request's output-token ceiling. A bounded 2026-08-23 strict JSON Schema smoke without an artificial low ceiling passed 7/8; Qwen plain chat and forced-tool calls passed, while the current strict JSON Schema shim still returned 503. This upstream smoke is not a paid Halo call, benchmark, or SLA.

claude-opus-5

High-end reasoning and coding workloads.

reasoningtoolsstructured

claude-sonnet-5

Balanced agent, chat, and code work.

reasoningtoolsstructured

deepseek-v4-flash

Fast coding, chat, and math requests.

fasttoolsstructured

deepseek-v4-pro

Reasoning-heavy code and math work.

reasoningtoolsstructured

glm-5.2

Reasoning and Chinese-language tasks.

chinesetoolsstructured

kimi-k2.7-code

Code generation and mathematical work.

codingtoolsstructured

minimax-m3

Chat and lightweight automation.

chattoolsstructured smoke

qwen3.5-397b-a17b

General reasoning, code, and agent work with at least a 128,000-token context window.

plain verifiedforced tool verifiedstrict schema degraded
Transparent AntSeed rates

Model-specific launch pricing.

USDC per one million tokens. Halo remains a flat 1 USDC per one million total tokens; compare against these AntSeed input/output rates for your traffic mix.

Current acquisition pricing · published by signed provider metadataCached input is 10% of input rate
ModelInput / 1MOutput / 1MCached input / 1M
claude-opus-50.160.490.016
claude-sonnet-50.090.290.009
deepseek-v4-flash0.0190.0390.0019
deepseek-v4-pro0.15840.31680.01584
glm-5.20.421.320.042
kimi-k2.7-code0.133680.62370.01338
minimax-m30.0090.0190.0009
Built for discovery

Agents should not have to read a sales page.

The machine catalog exposes models, payment rails, rates, capabilities, and live verification endpoints in JSON. Use the source directories when fresher network state matters.

agent discovery
// Start here
GET https://adfreellm.com/llm-services/catalog.json

// Then select a rail
{
  "halo": {
    "price_usdc_per_million_total_tokens": 1,
    "discovery": "https://relay.runhalo.xyz/v1/models"
  },
  "antseed": {
    "pricing": "per-model input/output",
    "metadata": "http://1.14.137.105/metadata"
  }
}
No monthly plan

Start with the rail your agent already understands.

Use Halo for one flat token price, or AntSeed when per-model economics matter. Both expose live public discovery before you commit funds.