336.AI
PromptsAPIPricing
PromptsAPIPricing
OverviewModel RankingAPI Key ManagementAPI DocsDev ToolsRequest LogsUsage AnalyticsBillingChat UsagePricingNotificationsTickets
Current Balance (Stars)
Top Up Now
3306.AI API
One-stop AI LLM API
3306.AI API
OverviewModel RankingAPI Key ManagementAPI DocsDev ToolsRequest LogsUsage AnalyticsBillingChat UsagePricingNotificationsTickets
Current Balance (Stars)
Top Up Now
3306.AI API
One-stop AI LLM API
API Platform
3306.AI DevelopersAPI Reference

API Documentation
OpenAI/v1Anthropic/anthropic

OpenAI- and Anthropic-compatible endpoints with ready-to-paste IDE integration guides.

Visible15
Compatibility
Models & billing
API capabilities
Account / Balance
Chat
Images
Audio
Embeddings
Models
Resources
System status
All systems operational
3306.AI APIOperational
OpenAI APIOperational
Claude APIOperational
API uptime
99.9%
Stars
Metered

Integration checklist

  • Authorization: Bearer <API_KEY>
  • One key works with both OpenAI and Anthropic protocols
  • Actual charges are shown in Usage and X-Stars-Cost
Card view

Pricing

The API bills per token. The table below lists the public rate for every callable model in US dollars, split into input, cached input and output, per 1M tokens; your balance is debited in Stars at 1 ★ = $0.01.

Billing unit

Rates are quoted in US dollars; balances settle in Stars, and one Star is fixed at $0.01. The Star figure returned with a response is exactly what your balance is debited, to four decimal places.

1 ★ = $0.01

Formula

Input tokens are split into a cache-miss part and a cache-hit part and priced separately at their dollar rates; output tokens are priced on their own. The three add up to the raw cost in dollars; dividing by the $0.01 Star value gives the Stars debited.

usd   = (input_miss / 1e6) * in_rate_usd
      + (input_cached / 1e6) * cache_rate_usd
      + (output / 1e6) * out_rate_usd

input_miss = max(0, prompt_tokens - cached_tokens)
stars      = usd / 0.01
charged    = max(1, round(stars, 4))   # 单位 ★

Every successful request carries a 1 Star minimum. If the raw cost lands below 1 Star you are charged 1 Star; above that you are charged the exact amount. A response with zero input and zero output tokens is not charged at all.

Public rates ($ / 1M tokens)

The model name is the exact string to put in the model field, identical to the id returned by GET /v1/models. Rates follow upstream changes; this page is the source of truth.

ModelDirect / Third-partyInputCached inputOutputLong context
Loading live rates…
The channel column says which tier the quoted rate comes from: official direct bills at the vendor list price, third-party providers discount that same list price by their own ratio. A model available on both tiers gets one row per tier; the third-party row carries its discount badge, and the struck-through figure under the discounted rate is the list price.
A dash in the cached-input column means the upstream publishes no separate cache-read rate for that model, so cached tokens are currently not billed separately. If a cache rate appears upstream, this page is updated with it.
Cache writes (Anthropic cache_creation_input_tokens) are billed at the input rate ×1.25 for the 5-minute TTL and ×2 for the 1-hour TTL (cache_control.ttl="1h"), matching the vendor’s published rates; subsequent hits are billed at the “Cache” rate.
Long context column: “> 200K ×2” means that once a request’s prompt tokens (cache hits included) exceed 200K, input, cache and output rates are all multiplied by 2. When routes for the same model differ, the cheapest tier is shown.

Image / Video / Audio (per call / per second)

Billed per call or per second. Final price = max(minimum charge, base × size multiplier × quality multiplier); each tier is pre-computed on the right.

ModelTypeDirect / Third-partyBaseSize / resolution tiersQuality tiersMin charge
Loading live rates…
Tier price = max(minimum charge, base × multiplier); when both size and quality are given the two multipliers stack. Failed generations are not charged.

Worked examples

Both token counts and Star amounts below are copied from real production calls, not estimated.

Example 1 · an ordinary chat turn

A ~64k-token long-context request where almost all of the input misses the cache.

model            deepseek-v4-flash
prompt_tokens    64012   (cached_tokens 2560)
completion       1

input_miss  61452 / 1e6 * $0.28    = $0.01720656
input_cached 2560 / 1e6 * $0.056   = $0.00014336
output          1 / 1e6 * $0.56    = $0.00000056
------------------------------------------------
usd                                  $0.01735
X-Stars-Cost   0.01735 / 0.01      = 1.735 ★
Example 2 · same prefix resent, cache hit

The identical prompt sent again straight away; upstream reports all 64k input tokens as cache hits.

model            deepseek-v4-flash   (same prompt, resent)
prompt_tokens    64012   (cached_tokens 64000)
completion       1

input_miss     12 / 1e6 * $0.28    = $0.00000336
input_cached 64000 / 1e6 * $0.056  = $0.00358400
output           1 / 1e6 * $0.56   = $0.00000056
------------------------------------------------
usd                                  $0.003588  ( = 0.3588 ★ )
X-Stars-Cost   max(1, 0.3588)      = 1 ★       ( = $0.01 )

The cache cuts the raw cost from 1.735 Stars to 0.3588 Stars — about 80% off — but that lands under the 1 Star minimum, so this call is charged 1 Star. The saving shows up on the bill once contexts are long and reuse is heavy.

Reading what a call cost

Every successful response carries X-Stars-Cost, the exact number of Stars debited, plus X-Request-Id for reconciliation and support tickets.

X-Stars-Cost: 1.735
X-Request-Id: req_f08c8151fc95438f9c5c524511a16806

On /v1/messages and /anthropic/v1/messages the cost is also written straight into the usage object:

"usage": {
  "input_tokens": 2203,
  "output_tokens": 1,
  "cost_stars": 1.5026,
  "cost_usd": 0.015
}

Billing rules

  • Only successful responses are billed. Auth failures, unknown models, invalid parameters and rate-limit rejections cost nothing.
  • Streaming is billed from the usage the upstream finally reports, at the same rates as non-streaming. A stream cut short is billed for the tokens actually produced.
  • A worst-case amount is held before the call and reconciled against real usage the moment upstream reports it, so nothing stays reserved after the request finishes.
  • An insufficient balance returns 402; there is no overdraft. Per-request detail is available on the usage page in the console, searchable by request id.