skip to content

Use your AI credits

Create a key, connect your client, and send your first request. The gateway supports OpenAI-style chat, image, and video requests.

Step 1: Create your API key

Open the dashboard, connect your wallet, and sign the sign-in message. Choose AI credits to earn credit on future splits, then create an API key.

Copy the full key when it appears. It starts with api_ and is shown only once. If you lose it, revoke it and create another. All keys for a wallet spend from the same credit balance.

You can create a key before you have credit. Paid requests need enough available credit to cover their reservation. Existing credits remain usable after you sell tokens or switch future rewards to SOL.

Step 2: Connect your client

In a tool that supports a custom OpenAI-compatible endpoint, set its base URL and API key to these values. The key comes from this site, not from an OpenAI account.

Base URLhttps://holdapi.lol/v1
API keyapi_…Your full key

Keep /v1 at the end of the base URL. Use a client’s Chat Completions mode. Features that require Responses, Assistants, embeddings, audio, or file uploads are not implemented by this gateway.

Keep the key in a local environment variable or a server secret. Do not put it in public browser code or a shared repository.

Step 3: Choose a model

Copy a full model ID from the catalog. IDs use the form provider/model. A display name such as “GPT” or “Claude” is not a model ID.

models · read from /v1/models

model id96 models
  • openai/gpt-6-astraopenai
  • openai/gpt-6.1-solopenai
  • openai/gpt-6-solopenai
  • openai/gpt-6-lunaopenai
  • openai/gpt-5.6-solopenai
  • openai/gpt-5.6-terraopenai
  • openai/gpt-5.6-lunaopenai
  • openai/gpt-5.6-sol-proopenai
  • openai/gpt-5.6-terra-proopenai
  • openai/gpt-5.6-luna-proopenai
  • openai/gpt-5.5openai
  • openai/gpt-5.5-proopenai
  • openai/chat-latestopenai
  • openai/gpt-5.4openai
  • openai/gpt-5.4-proopenai
  • openai/gpt-5.1openai
  • openai/gpt-5.2openai
  • openai/gpt-5.4-miniopenai
  • openai/gpt-5-miniopenai
  • openai/gpt-5.4-nanoopenai
  • openai/gpt-5.2-proopenai
  • openai/gpt-5.3-codexopenai
  • openai/gpt-4.1openai
  • openai/gpt-4.1-miniopenai
  • openai/gpt-4.1-nanoopenai
  • openai/gpt-4oopenai
  • openai/gpt-4o-miniopenai
  • openai/o1openai
  • openai/o3openai
  • openai/o3-miniopenai
  • openai/o4-miniopenai
  • anthropic/claude-haiku-4.5anthropic
  • anthropic/claude-sonnet-5.5anthropic
  • anthropic/claude-sonnet-5anthropic
  • anthropic/claude-sonnet-4.6anthropic
  • anthropic/claude-sonnet-4.5anthropic
  • anthropic/claude-opus-4.5anthropic
  • anthropic/claude-opus-4.7anthropic
  • anthropic/claude-fable-5.1anthropic
  • anthropic/claude-fable-5anthropic
  • anthropic/claude-opus-4.8anthropic
  • anthropic/claude-opus-5anthropic
  • anthropic/claude-opus-5.5anthropic
  • google/gemini-3.1-progoogle
  • google/gemini-3-flash-previewgoogle
  • google/gemini-3.8-flashgoogle
  • google/gemini-3.6-flashgoogle
  • google/gemini-3.5-flashgoogle
  • google/gemini-2.5-progoogle
  • google/gemini-2.5-flashgoogle
  • google/gemini-3.5-flash-litegoogle
  • google/gemini-3.1-flash-litegoogle
  • google/gemini-2.5-flash-litegoogle
  • deepseek/deepseek-v4-flash-vision-expdeepseek
  • deepseek/deepseek-v4-prodeepseek
  • deepseek/deepseek-chatdeepseek
  • deepseek/deepseek-reasonerdeepseek
  • moonshot/kimi-k3moonshot
  • zai/glm-5.3zai
  • zai/glm-5.3-flashzai
  • zai/glm-5.2zai
  • zai/glm-5.1zai
  • zai/glm-5zai
  • zai/glm-5-turbozai
  • xai/grok-4.3xai
  • xai/grok-build-0.1xai
  • xai/grok-4.7xai
  • xai/grok-4.6xai
  • xai/grok-4.5xai
  • xai/grok-imagine-videoxai
  • xai/grok-imagine-video-1.5xai
  • minimax/minimax-m2.7minimax
  • minimax/minimax-m3minimax
  • qwen/qwen3.8-maxqwen
  • qwen/qwen3.7-maxqwen
  • qwen/qwen3.7-plusqwen
  • qwen/qwen3.7-flashqwen
  • qwen/qwen3.8-flashqwen
  • tencent/hy4-previewtencent
  • xiaomi/mimo-v2.5xiaomi
  • xiaomi/mimo-v2.5-proxiaomi
  • nvidia/nemotron-3-nano-omni-30b-a3b-reasoningnvidia
  • nvidia/nemotron-3.5-lightningnvidia
  • nvidia/llama-3.2-11b-visionnvidia
  • nvidia/nemotron-3-ultra-550bnvidia
  • poolside/laguna-xs-2.1nvidia
  • nvidia/muse-glimmer-30bnvidia
  • cohere/north-mini-codecohere
  • poolside/laguna-s-2.1poolside
  • mistral/mistral-large-4mistral
  • bytedance/seedance-1.5-probytedance
  • bytedance/seedance-2.0-fastbytedance
  • bytedance/seedance-2.0-minibytedance
  • bytedance/seedance-2.0bytedance
  • bytedance/seedance-2.5bytedance
  • azure/sora-2azure

Use the same ID in your client’s model setting or the request’s model field. Choose a chat model for the chat example below, an image model for image generation, or a video model for the video section.

Step 4: Make your first request

On macOS, Linux, or a Bash shell, run the Settings snippet first. Replace api_your_dashboard_key with your full key and provider/model with a chat model ID from the catalog. Then copy the cURL, Python, or JavaScript example.

export OPENAI_BASE_URL="https://holdapi.lol/v1"
export OPENAI_API_KEY="api_your_dashboard_key"
export API_MODEL="provider/model"
openai compatible · Settings

The examples limit output to 256 tokens and turn off SDK retries. A smaller output limit can lower the reservation. If the model does not accept a parameter, use its supported equivalent.

On Windows PowerShell, set values with $env:OPENAI_API_KEY = 'your full key', and set OPENAI_BASE_URL and API_MODEL the same way. Then use the Python or JavaScript example.

Video generation

Video takes two requests, because a clip takes one to three minutes to render. The first starts the job and reserves its cost. It answers 202 with a job id and a poll_url. The second polls that job: 202 while it renders, then 200 with the clip URL in data[0].url.

curl "https://holdapi.lol/v1/videos/generations" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  --data @- <<EOF
{"model":"bytedance/seedance-1.5-pro","prompt":"a robot dog skating down a quiet street, slow dolly","duration_seconds":5,"resolution":"720p"}
EOF
openai compatible · Start the job

Video models are priced per second, not per token. The reservation covers the duration and resolution you ask for; duration_seconds and resolution default to the model’s own, and each model has a maximum duration. Read both from /models, where a video row carries type: "video" and pricing.usd_per_second.

Credits are charged on the poll that finds the job completed, and the reservation is released in full if the job fails. Polling costs nothing and is safe to repeat: a finished job returns the same clip URL and the same charge however many times you poll it.

Name a clip length as duration_seconds, in whole seconds. A field the gateway cannot price a video by is refused with a 400 naming it, rather than passed on where it could cost more than was reserved. If the provider quotes more than the reservation, the job is abandoned, the reservation is released, and the response is a 502 with both figures.

Only the key’s own wallet can poll a job. A job id from another wallet returns 404. Keep the id: without it, the reservation stays held until it is reconciled against the provider’s ledger.

API reference

Append these paths to the base URL. Every endpoint requires Authorization: Bearer api_… with your full key. POST bodies use Content-Type: application/json.

POST/chat/completions

Send a model ID and a messages array. Add stream: true for streaming. Optional parameters depend on the selected model.

POST/images/generations

Generate images with a supported image model and a prompt. Size and image count depend on that model.

POST/videos/generations

Start a video job with a video model and a prompt. Answers 202 with a job id and a poll_url. Optional duration_seconds and resolution depend on that model.

GET/videos/generations/{id}

Poll a video job. 202 while it runs, 200 with the clip URL when it is done. Polling is free, and only the key that started a job can poll it.

GET/models

List the model IDs available through the gateway. Use the full id value in your requests. Each row has a type: chat models are priced per million tokens, video models per second.

GET/balance

Read your wallet credit balance. credits_usd is the dollar balance; credits_micro_usd is the integer amount in millionths of a dollar.

Read your remaining credit

curl "$OPENAI_BASE_URL/balance" \
  -H "Authorization: Bearer $OPENAI_API_KEY"
openai compatible · cURL

How credits are spent

Before a request runs: the gateway reserves its estimated or quoted cost. Your available balance must cover the reservation, and the request must fit the per-call spending limit.

After the cost is known: usage-based billing settles the reservation to the provider’s reported cost and releases the difference. Quoted, prepaid requests are charged at the signed quote instead. Model pricing and the payment method affect the final cost.

Streaming responses and interrupted requests can leave a reservation pending until the provider’s cost is reconciled. A confirmed uncharged request releases its hold; a timeout alone does not confirm that a request was free.

If the final charge exceeds the reservation and your available balance, the unpaid amount is recorded. Future credit rewards first repay that amount before increasing your spendable balance. SOL rewards are not used to repay it.

Response headers include x-api-cost-micro-usd, x-api-reserved-micro-usd, and x-api-credits-remaining-micro-usd. For your current balance after a stream completes, read /balance or the dashboard.

One dollar equals 1,000,000 micro-USD. Creating another key does not create another balance or bypass the wallet’s limits.

Troubleshooting

Errors use an OpenAI-style JSON body with an error object. Read its code and message to distinguish errors with the same HTTP status.

400

Invalid request or model

Check the JSON body and use a full provider/model ID from the catalog. Choose the route for the model type.

401

Invalid API key

Use a complete, unrevoked api_ key from your dashboard. Send it as Authorization: Bearer followed by the key.

402

Insufficient credits

The balance must cover the reservation before a request starts. Check /balance, reduce max_tokens or the image count, or wait for more credits.

402

Request exceeds the spending limit

quote_over_limit means the estimate or quote is above the per-call ceiling. Reduce the requested output or image count; more balance alone does not lift this limit.

429

Too many requests

Slow down or reduce concurrent requests. Rate and concurrency limits are shared across your wallet’s keys.

502 / 503

Upstream service unavailable

Check your balance before retrying an uncertain request. A missing answer does not prove that the provider did not bill it.

If an answer is lost after a request is paid, retrying can create a second paid request. Check your balance and avoid blindly retrying an uncertain result.