Skip to main content

api reference

Base URL · https://www.lazysusan.ai/api/v1

LazySusan API

One key and one OpenAI-compatible endpoint for chat, image and video models from every major lab. Change one string to switch models.

Overview

The LazySusan API gives you 18 chat models, 13 image models and 6 video models behind one base URL and one key. Requests and responses follow the OpenAI API format, so the official OpenAI SDKs and most tools built on them work after you change the base URL and key.

EndpointWhat it does
POST /chat/completionsChat with any chat model. Streaming, images in, function calling and JSON output.
POST /images/generationsGenerate images. Returns hosted URLs or base64.
POST /videosStart a video. It runs in the background; poll for the result.
GET /videos/{id}Check a video and get its URL when it's done.
GET /videosYour 50 most recent videos.
GET /modelsEvery model with its limits and price. No key needed.
GET /creditsYour remaining balance.

You pay from a prepaid balance of credits. One credit is $0.001, so $25 buys 25,000 credits. Every request is billed for what it actually used, and requests that fail aren't charged.

Quickstart

  1. Sign in and open the developer console.
  2. Buy credits (packs start at $25).
  3. Create a key. Copy it now: it's shown once.
  4. Send your first request:
curl https://www.lazysusan.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $LAZYSUSAN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-5-5",
    "messages": [{"role": "user", "content": "Write a haiku about lazy susans."}]
  }'

If you already use the OpenAI SDK, the only changes are base_url and api_key. The response includes usage.cost_credits, the exact amount this request cost.

Authentication

Send your key on every request as Authorization: Bearer lsk_live_…. The x-api-key: lsk_live_… header works too. Keys are lsk_live_ followed by 40 letters and digits.

  • Keep keys on your server. Don't put them in browser or mobile app code, and don't commit them to git.
  • Create a separate key per app or environment, so you can revoke one without breaking the others. You can have up to 20 active keys.
  • Revoking a key in the console takes effect on the next request.
  • We store only a one-way hash of each key, so a lost key can't be shown again. Create a new one.

Models and prices

Model ids are provider/model, for example openai/gpt-6.1-sol or google/veo-3.1-fast. Prices are in credits ($0.001 each). The same list, with every limit, is available from GET https://www.lazysusan.ai/api/v1/models (no key needed) and GET /models/{provider}/{model} for a single model.

Chat models

Priced per million tokens. Output includes any reasoning (“thinking”) tokens the model uses.

Chat models and prices
Model idContextMax outputInput / 1MOutput / 1MImages in
openai/gpt-6.1-sol400K64,0002,30011,500yes
openai/gpt-6-astra400K64,00011,50057,500yes
openai/gpt-6-luna400K64,000115575yes
anthropic/claude-fable-5-11,000K64,00011,50057,500yes
anthropic/claude-opus-5-51,000K64,0004,60023,000yes
anthropic/claude-sonnet-5-51,000K64,0002,30011,500yes
anthropic/claude-haiku-5-51,000K64,000115575yes
google/gemini-3.1-pro1,000K65,5362,30013,800yes
google/gemini-3.8-flash1,000K65,536862.54,312.5yes
google/gemini-3.5-flash-lite1,000K65,5363452,875yes
xai/grok-4.71,000K32,7682,3006,900yes
xai/grok-4.31,000K32,7681,437.52,875yes
xai/grok-build-0.11,000K32,7681,1502,300yes
deepseek/deepseek-flash128K8,1923451,380no
deepseek/deepseek-v4-pro128K8,1921,5184,554no
perplexity/sonar127K8,1922,3002,300no
perplexity/sonar-pro127K8,1926,90034,500no
perplexity/sonar-reasoning-pro127K8,1924,60018,400no

Cached input (where the provider caches a repeated prompt prefix) is billed at the lower cached rate shown in GET /models. Long prompts on some models are billed at a higher tier by the provider: Claude Haiku 5.5 above 100K input tokens, Gemini 3.1 Pro above 200K and Grok above 200K. Perplexity Sonar models run a live web search on every request; the search adds a per-request fee and search-result tokens, which are included in the request's cost_credits.

Image models

Image models and prices
Model idSizesPrice per imageMax nInput images
google/nano-banana-2.11024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x124838.644yes
google/nano-banana-pro1024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x1248154.14yes
openai/gpt-image-2.51024x1024, 1536x1024, 1024x1536by usage: typically 5.635 (low), 19.435 (medium), 48.415 (high)4no
openai/gpt-image-2.5-fast1024x1024, 1536x1024, 1024x1536by usage: typically 5.635 (low), 19.435 (medium), 48.415 (high)4no
bfl/flux-2-pro1024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x124857.51no
bfl/flux-31024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x124855.21no
bfl/flux-kontext-pro1024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x1248461no
bfl/flux-kontext-max1024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x1248921no
bfl/flux-pro-1.11024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x1248461no
bfl/flux-pro-1.1-ultra1024x1024, 1344x768, 768x1344, 1184x864, 864x1184, 1248x832, 832x1248691no
xai/grok-imagine-image1024x1024234no
xai/grok-imagine-image-quality1024x102457.54no
xai/grok-imagine-image-2.01024x1024694no

GPT Image models are billed by the tokens each image actually uses (34,500 credits per 1M image output tokens, 5,750 per 1M prompt tokens), so the figures above are typical, not fixed. Before the call, a request reserves the most an image of that size and quality could use and is charged only what it used. Larger sizes and higher quality use more tokens.

Video models

Video models and prices
Model idSecondsResolutionAspect ratioPrice per videoNotes
google/veo-3.14, 6, 8720p, 1080p16:9, 9:16720p · 4s: 1,840; 720p · 6s: 2,760; 720p · 8s: 3,680; 1080p · 8s: 3,680includes audio
google/veo-3.1-fast4, 6, 8720p, 1080p16:9, 9:16720p · 4s: 460; 720p · 6s: 690; 720p · 8s: 920; 1080p · 8s: 1,104includes audio
google/veo-3.1-lite4, 6, 8720p, 1080p16:9, 9:16720p · 4s: 230; 720p · 6s: 345; 720p · 8s: 460; 1080p · 8s: 736includes audio
runway/gen-4.55, 10720p16:9, 9:16720p · 5s: 690; 720p · 10s: 1,380—
runway/gen-4-turbo5, 10720p16:9, 9:16720p · 5s: 287.5; 720p · 10s: 575needs a starting image
minimax/hailuo-2.36, 10768p, 1080p16:9768p · 6s: 322; 768p · 10s: 644; 1080p · 6s: 563.5—

Switching models

Every model in a category takes the same request, so switching is a one-string change. No new SDK, key or billing account:

for model in ["openai/gpt-6.1-sol", "anthropic/claude-opus-5-5", "google/gemini-3.8-flash", "xai/grok-4.7"]:
    reply = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": "Summarise our refund policy in one sentence."}],
        max_tokens=400,
    )
    print(model, reply.usage.cost_credits, reply.choices[0].message.content)

Chat completions

POST https://www.lazysusan.ai/api/v1/chat/completions

Chat completion parameters
ParameterTypeNotes
modelstring, requiredA chat model id from the table above.
messagesarray, required1–2,000 messages. Roles: system, developer, user, assistant, tool. content is a string or an array of parts ({type:"text",text} and {type:"image_url",image_url:{url,detail?}}). Assistant messages may have content null with tool_calls. Tool messages need tool_call_id. An optional name is kept (up to 64 characters).
max_tokensintegerMost tokens to generate, including reasoning tokens. max_completion_tokens is accepted as an alias. Up to the model's max output. Default 4,096; if your balance can't cover that and you didn't set it, the request runs with 1,024.
temperaturenumber 0–2Ignored by OpenAI GPT-6 models.
top_pnumber 0–1Ignored by OpenAI GPT-6 models.
presence_penaltynumber −2–2Ignored by OpenAI GPT-6 models.
frequency_penaltynumber −2–2Ignored by OpenAI GPT-6 models.
seedintegerBest-effort determinism where the model supports it.
stopstring or string[]Up to 4 stop sequences.
toolsarrayUp to 128 function tools: {type:"function", function:{name, description?, parameters?}}. Other tool types (for example built-in web search) are rejected.
tool_choicestring or objectPassed to the model as given.
parallel_tool_callsboolean
response_formatobject{type: "text" | "json_object" | "json_schema", …}.
streambooleanStream tokens as server-sent events. See Streaming.
stream_optionsobject{include_usage: true} adds a final chunk with usage and cost.
nintegerOnly 1 is supported; any other value is rejected.

Other parameters are ignored rather than rejected, so code written for other OpenAI-compatible APIs keeps working. Reasoning models (GPT-6, Gemini 3.x, Grok, DeepSeek V4 Pro, Sonar Reasoning Pro) spend part of max_tokens thinking before they answer. With a very small limit they can return an empty answer with finish_reason: "length": leave room, a few hundred tokens or more.

Response

200 OK · application/json
{
  "id": "gen-6f0c2a5e-1d2b-4c7e-9a51-0b8f3c2d7e14",
  "object": "chat.completion",
  "created": 1791676800,
  "model": "anthropic/claude-sonnet-5-5",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Turns once, then rests…" },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 18,
    "completion_tokens": 22,
    "total_tokens": 40,
    "cost_credits": 0.295
  }
}

usage.completion_tokens includes reasoning tokens. usage.prompt_tokens_details.cached_tokens appears when part of the prompt was cached. Perplexity models also return citations and search_results. Every response carries an x-request-id header with the same id as the body. Include it when you contact support.

Streaming

With "stream": true the response is text/event-stream. Each event is a data: line holding a chat.completion.chunk, and the stream ends with data: [DONE]. With "stream_options": {"include_usage": true}, one more chunk arrives just before [DONE], with an empty choices array and the request's usage, including cost_credits.

data: {"id":"gen-…","object":"chat.completion.chunk","model":"openai/gpt-6-luna","choices":[{"index":0,"delta":{"content":"1, 2"}}]}

data: {"id":"gen-…","object":"chat.completion.chunk","model":"openai/gpt-6-luna","choices":[{"index":0,"delta":{"content":", 3"},"finish_reason":"stop"}]}

data: {"id":"gen-…","object":"chat.completion.chunk","model":"openai/gpt-6-luna","choices":[],"usage":{"prompt_tokens":14,"completion_tokens":9,"total_tokens":23,"cost_credits":0.029}}

data: [DONE]

If you disconnect mid-stream, the request stops and you're billed for the prompt and the output generated up to that point. If a provider fails mid-stream, you'll get a final event of the form {"error": {"message", "type": "upstream_error"}} and the stream closes without [DONE].

Images in, tools and JSON

Images in a prompt

Models with “Images in: yes” accept up to 20 images per request as image_url parts. Use a public https:// URL or a base64 data:image/(png|jpeg|webp|gif);base64,… URL. detail may be low, high or auto.

client.chat.completions.create(
    model="google/gemini-3.8-flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What's in this photo?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
        ],
    }],
)

Function calling

tools = [{
    "type": "function",
    "function": {
        "name": "get_order_status",
        "description": "Look up an order by id",
        "parameters": {
            "type": "object",
            "properties": {"order_id": {"type": "string"}},
            "required": ["order_id"],
        },
    },
}]
first = client.chat.completions.create(
    model="openai/gpt-6.1-sol",
    messages=[{"role": "user", "content": "Where is order A-1042?"}],
    tools=tools,
)
call = first.choices[0].message.tool_calls[0]
# …run your function, then send the result back:
second = client.chat.completions.create(
    model="openai/gpt-6.1-sol",
    messages=[
        {"role": "user", "content": "Where is order A-1042?"},
        first.choices[0].message,
        {"role": "tool", "tool_call_id": call.id, "content": '{"status": "shipped"}'},
    ],
    tools=tools,
)

JSON output

Set response_format to {"type": "json_object"}, or json_schema with a schema, on models that support it. If a model doesn't, the provider rejects the request with a 400 and you aren't charged.

Images

POST https://www.lazysusan.ai/api/v1/images/generations

Image generation parameters
ParameterTypeNotes
modelstring, requiredAn image model id.
promptstring, requiredUp to 4,000 characters.
nintegerDefault 1, up to the model's Max n (FLUX: 1; Nano Banana, GPT Image and Grok Imagine: 4).
sizestringOne of the model's sizes, e.g. 1024x1024 (the default) or 1344x768. Nano Banana and FLUX map sizes to aspect ratios.
qualitystringGPT Image only: low, medium (default) or high. auto and standard mean medium; hd means high.
response_formatstringurl (default): a permanent hosted URL. b64_json: the image inline as base64.
imagestringNano Banana models only: an image to edit or use as a reference (https URL or data: URL, PNG/JPEG/WebP, up to 20 MB). Or send up to 3 as images: [].
curl https://www.lazysusan.ai/api/v1/images/generations \
  -H "Authorization: Bearer $LAZYSUSAN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "google/nano-banana-2.1", "prompt": "A red ceramic cube on a white table, studio light", "size": "1344x768"}'
200 OK · application/json
{
  "id": "img-2b7d…",
  "object": "image.generation",
  "created": 1791676800,
  "model": "google/nano-banana-2.1",
  "data": [{ "url": "https://…/api/…/2b7d…-0.png" }],
  "usage": { "images": 1, "cost_credits": 38.64 }
}

You're charged only for images actually delivered: if you ask for 4 and the provider returns 3, you pay for 3. If no image comes back (for example the prompt was declined by the provider's safety filter) the request returns 400 and costs nothing. Hosted URLs stay up and are unguessable; anyone with the URL can open it, so treat it like a share link. In the rare case hosting is unavailable, an item comes back as b64_json even if you asked for url.

Videos

Video takes from about a minute to a few minutes, so it runs in the background. POST https://www.lazysusan.ai/api/v1/videos returns 202 Accepted straight away with the video's id and status queued; then poll GET https://www.lazysusan.ai/api/v1/videos/{id} every 10 seconds or so until the status is succeeded or failed.

Video parameters
ParameterTypeNotes
modelstring, requiredA video model id.
promptstring, requiredUp to 2,000 characters.
secondsintegerOne of the model's durations. Default: the shortest.
resolutionstringOne of the model's resolutions. Default: the first listed. Veo renders 1080p only at 8 seconds.
aspect_ratiostring16:9 (default) or 9:16.
imagestringStarting frame (https or data: URL, PNG/JPEG/WebP, up to 20 MB). Required for runway/gen-4-turbo.
negative_promptstringGoogle Veo models only. Up to 1,000 characters.
import time, requests

H = {"Authorization": f"Bearer {os.environ['LAZYSUSAN_API_KEY']}"}
job = requests.post(f"https://www.lazysusan.ai/api/v1/videos", headers=H, json={
    "model": "google/veo-3.1-fast",
    "prompt": "A slow pan across a calm lake at sunrise",
    "seconds": 8,
    "resolution": "1080p",
}).json()

while job["status"] in ("queued", "running"):
    time.sleep(10)
    job = requests.get(f"https://www.lazysusan.ai/api/v1/videos/{job['id']}", headers=H).json()

print(job["status"], job.get("video_url"), job.get("cost_credits"))
GET /videos/{id} · succeeded
{
  "id": "vid-91c4…",
  "object": "video",
  "model": "google/veo-3.1-fast",
  "status": "succeeded",
  "created": 1791676800,
  "completed": 1791676911,
  "seconds": 8,
  "resolution": "1080p",
  "aspect_ratio": "16:9",
  "video_url": "https://…/api/…/91c4…-0.mp4",
  "cost_credits": 1104
}
  • The POST response includes cost_credits_held: the video's exact price, held from your balance while it renders.
  • On success you're charged that amount. On failure, the status is failed, error.message says why, and the hold is returned in full (cost_credits: 0).
  • A video that hasn't finished after 2 hours is marked failed and refunded.
  • Videos finish even if you stop polling; GET /videos lists your 50 most recent. The video URL is permanent and unguessable.

Credits and billing

The API runs on a prepaid balance. One credit is $0.001. Nothing is billed in arrears: a request only runs if your balance can cover it.

Buying credits

PackCredits
$2525,000
$100100,000
$500500,000
$2,0002,000,000

Buy packs in the console with any card.

Automatic top-up

Turn on auto top-up to never run dry: when your balance falls below a threshold you choose (1,000 to 10,000,000 credits), we charge the card you last paid with for an amount you choose ($25 to $5,000) and add the credits straight away. It needs one manual purchase first, so we have a card on file. After 3 declined charges in a row, auto top-up switches itself off.

Monthly spend limit

Set a monthly limit in credits to cap spend. When a request would take the month past it, it's refused with 402 and the code monthly_limit_reached. Credits held by requests in progress count toward the limit.

How a request is billed

  1. Before a request runs, we hold the most it could cost: the full prompt plus max_tokens of output for chat, the per-image price × n for images, the clip price for video.
  2. We run it, and the model reports exactly what was used.
  3. You're charged for actual usage (never more than the hold), and the rest of the hold is released immediately.

Holds mean many requests can run at once without ever overdrawing your balance. A large max_tokens on an expensive model needs a bigger balance to start, even though you're only charged for what's used. Requests rejected before they reach a model, and requests the provider fails, cost nothing.

Checking your balance

curl https://www.lazysusan.ai/api/v1/credits -H "Authorization: Bearer $LAZYSUSAN_API_KEY"

{"object":"credits","balance":24183.552,"currency_per_credit_usd":0.001,"auto_top_up":false}

The console shows every request with its model, tokens and cost, plus spend this month by model.

Errors

Errors use the OpenAI shape, with an HTTP status and a machine-readable code:

402 Payment Required
{
  "error": {
    "message": "Not enough credits. This request can cost up to 62.445 credits. Add credits at https://www.lazysusan.ai/developers/console or lower max_tokens.",
    "type": "insufficient_balance",
    "code": "insufficient_balance",
    "param": null
  }
}
Error types
StatusTypeCommon codesWhat to do
400invalid_request_errorinvalid_json, model_not_found, invalid_request, images_not_supported, invalid_image, image_unreachable, image_too_large, image_required, context_length_exceeded, provider_rejected_request, no_imageFix the request; param names the field. provider_rejected_request means the model refused it (for example a safety filter); you weren't charged.
401authentication_errormissing_api_key, invalid_api_keyCheck the key and the Authorization header.
402insufficient_balanceinsufficient_balance, monthly_limit_reachedAdd credits, lower max_tokens, or raise the monthly limit.
403permission_erroraccount_suspendedContact support.
404not_found_errormodel_not_found, video_not_foundCheck the id.
405invalid_request_errormethod_not_allowedUse POST for chat.
413invalid_request_errorrequest_too_largeSend a smaller body (see Limits).
429rate_limit_errorrate_limitedWait for the Retry-After seconds, then retry.
500server_errormodel_unavailableRetry; if it persists, contact support with the request id.
502upstream_errorprovider_error, provider_unreachableThe model provider failed. Retry, or switch models. You weren't charged.
503upstream_errorprovider_rate_limitedThe provider is busy. Retry after Retry-After.

Responses include an x-request-id header (gen-…, img-… or vid-…) whenever a request got far enough to be billed. Send it to support if anything looks wrong.

Rate limits and limits

LimitValue
Requests per key120 per minute
Requests per account600 per minute, across all keys
Chat request body8 MB
Image request body64 MB
Video request body32 MB
Messages per chat request2,000
Images in a chat request20
Input images for image models3, each up to 20 MB
Image prompt4,000 characters
Video prompt2,000 characters
Active keys per account20

Over the rate limit you get 429 with a Retry-After header. Need more throughput? Write to support with your expected volume.

Data and security

  • All traffic is HTTPS. Keys are stored only as one-way hashes.
  • We don't store your prompts or the text models send back. We keep a usage record per request (model, token counts, cost, time) for billing and your console.
  • Generated images and videos are hosted at unguessable URLs so you can use them directly. Download and re-host them if you need to control access.
  • Image URLs you send us are fetched only from the public internet; private and internal addresses are refused.
  • Requests are processed by the model's provider under that provider's API terms.

Changelog

DateChange
2026-10-11Launch: 18 chat models, 13 image models and 6 video models; prepaid credits, auto top-up and monthly limits.

Support

Email support@lazysusan.ai with your request id. For volume pricing, invoicing or a higher rate limit, say so in the subject .