NanoGPT help
Ask about NanoGPT pages, models, API, memory, media generation, pricing, and support.
Press Enter to send. Press Shift+Enter for a new line. 0/2000 characters.
When an AI-generated answer is needed, your question and recent help-chat context may be sent to an external AI model provider for processing. The transcript stays in this browser tab for the session. Do not paste API keys, passwords, payment details, or other secrets. This assistant is limited to NanoGPT help; use normal chat for general AI tasks. Read the privacy policy.
Frequently asked questions
Browse common answers or ask the assistant above for a more specific path.
NanoGPT generally approves good-faith refund requests, subject to its Refund Policy. Examples include duplicate charges, clear billing errors, and some requests near the start of a subscription period. Extensive completed usage, abuse, fraud risk, or legal and payment-processor restrictions can affect eligibility.
Canceling a subscription stops future renewals; it does not automatically refund the current period. To request a refund, contact support@nano-gpt.com with your Support Key, account email, and reason, or open a private support ticket from the affected account. Include the payment reference and relevant screenshots, but never full card details or private wallet credentials.
The Help assistant cannot inspect payments or approve refunds. Support reviews the request, and applicable non-waivable consumer rights still apply.
Open Settings → Account → Delete account, review the permanent-deletion warning, and type DELETE to confirm. Export anything you need first. The deletion flow also clears local chat data in the current browser.
For individual chats, cloud snapshots, shared conversations, or API request logs, use their own deletion controls. Local copies on other devices and downloaded exports need to be removed separately. Some information may be retained to comply with legal obligations; deleting an account is not a promise that every provider erases data already sent to it.
If you cannot access the affected account or need help with a specific deletion request, open a private support ticket. The Help assistant cannot delete data or verify account state for you.
Check the generation in Media and its entry in Usage. A pending job or missing preview does not by itself establish that the provider failed. Avoid repeatedly submitting the same paid generation while its status is unclear.
Usage shows available request details and refund records. If the job failed, the result is missing, or the charge still looks wrong, open a private support ticket with the model, approximate time, amount, job or request ID, and error message or screenshot. Include the X-Request-ID response header for API calls when available.
The Help assistant cannot inspect your charge, issue a refund, or promise an outcome. Support can investigate the specific generation and billing records.
In chat, enable Private Mode for a private-capable model. For compatible API clients, use the NanoGPT Private Mode local proxy and its setup guide. A normal request to a TEE model is not automatically a Private Mode request.
Private Mode verifies the Tinfoil enclave and encrypts the request in your browser or local proxy before NanoGPT receives its body. NanoGPT still sees account, routing, timing, size, status, and usage metadata. Failed verification does not silently fall back to an ordinary plaintext route.
Chat history stays local by default. Syncing Private Mode conversations requires passphrase or passkey end-to-end encryption; ordinary cloud sync is blocked for those messages. Use the privacy guide and verification page for current setup and limits.
Choose the endpoint that matches your client: POST /api/v1/chat/completions for OpenAI Chat Completions clients, POST /api/v1/responses for Responses clients, or POST /api/v1/messages for Anthropic Messages clients. Their request bodies and streaming events differ; changing the URL alone does not convert one format to another.
Use the endpoint documentation and the dedicated integration guide for your client. Check the selected model and endpoint for tool, media, reasoning, and structured-output support. Account billing and API-key permissions still apply.
NanoGPT uses account funds for model usage. Add funds from Balance with credit card or supported crypto, then use chat, media generation, API, or other tools.
You can start with an anonymous local session. Creating an account is optional, but it makes balances easier to keep across devices.
Open a new chat, choose a text model or Auto Model, and type your message in the input box. Conversations are separate by default, although Context Memory and Global Memory can carry selected context forward when enabled.
No. NanoGPT can be used with an anonymous local session. Accounts are useful when you want a balance that is easier to keep across devices.
Chat history stays in your browser by default. Optional cloud sync is configured in Settings → Storage. Signing in alone does not restore unsynced chats.
NanoGPT can be installed as a Progressive Web App. The NanoGPT Browser Assistant is available for Chrome and Firefox, letting you attach pages, elements, images, or marked areas and then summarize, explain, translate, or chat with that context while you browse.
A model is the AI system that answers, reasons, writes code, analyzes images, or generates media. NanoGPT lets you choose between many text, image, video, and audio models.
Different models have different strengths, speeds, context sizes, prices, and media capabilities.
Some models are strongest at reasoning, some at coding, some at long context, some at vision, and some at low-cost high-volume work. Price usually reflects capability, provider cost, and performance.
Model pages show capabilities and pricing. Leaderboards help compare model quality across text, image, and video tasks.
Auto Model lets NanoGPT select a suitable Basic, Standard, or Premium model tier for the task. It is useful when you care more about the result than manually choosing a specific model.
Yes. Supported models and providers can reuse repeated prompt prefixes to lower latency and input costs. Implicit caching is automatic on eligible routes; keep reusable instructions and reference material at the start of the prompt.
Claude supports explicit cache controls such as `prompt_caching` or inline `cache_control`. For supported models, `caching: true` selects a cache-capable provider; it does not itself add explicit cache-write annotations. Follow the prompt-caching documentation for the selected model and provider.
Cache hits can lower pay-as-you-go cost, but cached input tokens still count toward your subscription input-token allowance.
Vision means a text model can accept images or screenshots alongside text. Vision models can describe images, read text, analyze screenshots, and reason about visual content.
NanoGPT connects supported models to external AI model providers. Availability and routing depend on the selected model and may change over time.
Check the model page or provider-selection controls for the options available for that model.
Ask the Help assistant with the task and required inputs. It can compare published model capabilities and context limits, then point to suitable model detail pages. Use those pages to compare current pricing and availability. If you prefer an automatic choice, select Auto Model and choose the Basic, Standard, or Premium tier that fits the task.
For manual selection, check the capabilities that matter: coding agents often need tool calling and enough context for repository files; screenshots need vision support, while PDFs can use native PDF input or text extraction in chat; and roleplay or translation requests should name the desired language, tone, and content requirements.
Model availability and capabilities change, so verify the recommendation on the current model detail page before relying on it.
Not every model exposes reasoning text, and some third-party clients ignore separate reasoning fields even when the model returns them. The normal API uses the `reasoning` field, while older clients may expect `reasoning_content`.
If you can change the request, set `reasoning.delta_field` to `reasoning_content`, or use the `reasoning_delta_field` or `reasoning_content_compat` shorthand. If you cannot change the request payload, use `https://nano-gpt.com/api/v1legacy/chat/completions` for the legacy field name.
Use `https://nano-gpt.com/api/v1thinking/chat/completions` when a client ignores reasoning-specific fields and should receive reasoning in the normal content stream. This is the preferred reasoning-compatible endpoint for JanitorAI.
Use the Media page in image mode. Enter a prompt, choose an image model, adjust settings such as aspect ratio or resolution when available, and start generation.
The help assistant can also recommend an image model for a specific prompt and link you directly to the right Media page setup.
Yes. Some image models support image-to-image or editing workflows. Choose a compatible model from the image generation flow and upload the source image where the UI offers image input.
As between you and NanoGPT, and to the extent permitted by applicable law, NanoGPT assigns its rights in generated output to you. Generated images may be used commercially, subject to applicable law and any restrictions in the selected model provider's terms.
You are responsible for checking third-party rights and provider-specific restrictions. Generated output is not guaranteed to be copyrightable, exclusive, or free of third-party rights.
Use the Media page in video mode. Choose a video model, enter the prompt, and adjust model-specific settings such as duration, aspect ratio, resolution, or image input when supported.
Some video models support image-to-video workflows for animating an existing image.
Image and video settings depend on the selected model. The Media page shows available settings such as aspect ratio, resolution, duration, rendering speed, and model-specific options.
NanoGPT has Context Memory and Global Memory features for carrying useful context forward. Memory can be controlled through conversation and system prompt settings.
Use the memory documentation for the exact setup and behavior.
Open Conversations in the browser where you created the chat. History stays on that device by default. If you enabled cloud sync in Settings → Storage, synced chats can also be restored after signing in on another device. Signing in alone does not restore unsynced chats.
Conversation history stays in the browser on the device where it was created by default. Open Conversations on that device to find and continue old chats.
Optional cloud sync can keep selected user data and conversations available across signed-in devices. The Storage tab in Settings controls NanoGPT-hosted sync or your own remote storage.
Unsynced local chats are not automatically recoverable in another browser. A synced chat can be restored after sign-in; an exported backup must be imported manually.
The Workspace groups related chats, files, notes, and tasks into projects. Create or open a project from Workspace, then start project-linked conversations so the project context and selected files stay organized together.
Projects are separate from ordinary chat history. Use Workspace for ongoing work with shared project context, and Conversations for the full list of individual chats.
Attach supported files from the chat input or add reusable files to a Workspace project. Image and other media capabilities vary by model. PDFs can also be used with models that do not have native PDF input: chat extracts the document text before sending it to the model.
If a model cannot read an attachment, try a model whose detail page lists the required capability, reduce an oversized file, or add the extracted text directly. Project files can be reused by project-linked conversations.
Yes, in chat. The Native PDF input badge means that the model can receive the PDF file directly. Without it, NanoGPT extracts text from uploaded PDFs and sends that text to the model. Scanned pages may need OCR; follow the prompt if OCR requires consent or payment.
Text extraction and OCR may lose images, charts, and page layout. For questions about visual details, choose a model with native PDF input or upload the relevant pages as images to a vision model.
The API does not automatically apply the chat upload extraction workflow. For models without native PDF input, extract the PDF text yourself and send it in messages, or render pages as images for a vision model. Kimi K3 supports document questions through extracted text; its vision capability does not imply direct PDF input support.
Export conversations from the Conversations page and keep the backup somewhere safe; restoring an export requires importing it manually. Optional cloud sync in Settings → Storage can restore synced chats across signed-in devices. Unsynced chats are not recovered by signing in alone.
Creating an account is the easiest way to keep balances across devices. Without an account, your local session identifier ties your browser to your balance; keep it private and do not clear your only session before securing access.
NanoGPT is designed so you can use the site without creating an account. Conversations are stored locally by default.
When you send a message to a model, the prompt and relevant conversation context are sent to the selected model provider so the model can answer. Provider-specific privacy and retention policies may apply.
ZDR is a best-effort routing commitment, not an independently verified guarantee about a provider’s servers. NanoGPT does not run the model inference itself. When ZDR-only routing applies, we route only through providers and routes that state they use Zero Data Retention, but we cannot inspect or audit their internal infrastructure and ultimately rely on them to honor that claim.
NanoGPT can guarantee only the behavior of systems it controls: by default, we do not log or retain prompt and response content as part of normal inference. We still keep limited non-content metadata needed for billing, reliability, and abuse prevention. Separate API request/response logging is off by default and stores content only when the API-key owner explicitly enables it.
NanoGPT does not attach IP addresses to prompts, model-provider requests, usage records, support tickets, or bug reports. Raw IP addresses are used temporarily in rate-limit and abuse-prevention systems and are automatically deleted after the relevant security window expires.
Retention varies by control, and NanoGPT does not publish individual thresholds because doing so could weaken those protections. Password-reset abuse protection may also store a keyed one-way identifier derived from an IP address as a separate security record.
NanoGPT does not intentionally write raw IP addresses into application log messages. Vercel separately processes network information as the hosting provider, and its public privacy policy does not give one universal IP-specific retention period that NanoGPT can promise on Vercel’s behalf.
Web search is separate from model inference. The selected search provider handles the search query. You can change web search settings when the feature is available for your workflow.
When Linkup is used through NanoGPT with our managed credentials, it operates under Zero Data Retention (ZDR). This applies both to Linkup web search and to the Linkup Research models. If you use your own Linkup API key, the data terms for your Linkup account apply instead.
TEE model protections apply to model execution and do not automatically extend to web-search provider requests.
Open Balance and choose a payment method. NanoGPT supports credit card payments and supported crypto deposits.
Crypto deposits are credited based on the current USD exchange rate. Nano deposits are converted to USD at the live rate, and new Nano deposits are no longer held as a Nano balance.
Every crypto and stablecoin top-up gets a 3% bonus: we credit 3% more than you pay. Card payments are credited at face value.
NanoGPT is pay-as-you-go by default. Text models are generally priced by input and output tokens. Media models are priced by model-specific generation settings such as resolution, duration, or rendering mode.
Pricing pages and model detail pages show the relevant costs.
The Pro plan currently advertises 60 million included input-token units per week and 100 included images per day. Cached input tokens count toward the included input-token allowance: cache hits can lower pay-as-you-go cost, but they do not reduce subscription quota usage. Most included text models consume one allowance unit per input token, while models marked with a 2x multiplier consume two. These figures can change, so check the live Subscription page before purchasing or planning usage. The weekly allowance resets Monday at 00:00 UTC; purchasing, renewing, resuming, or being billed does not reset it.
The included model list can change. Use the Subscription page for the current text and image model lists. Video generation, voice, web search, extended memory, TEE variants, and other paid extras are not included unless the live Subscription page explicitly says otherwise.
If an included limit is reached, included usage pauses until reset unless “Use balance after limits” is enabled, in which case eligible overage can use the normal pay-as-you-go balance.
For API overage, the API key must also allow balance spending: its billing mode must not be “Subscription only” and it must not have a $0 spend limit. Otherwise the request continues to return 429 until the included limit resets.
Yes. Website and API requests share the same subscription allowance. Use the subscription endpoint and an included model when you want the request covered by the subscription.
Use `GET /api/subscription/v1/models?detailed=true` for the current subscription-included text models and `POST /api/subscription/v1/chat/completions` for subscription-covered chat requests. Explicit provider selection is always pay-as-you-go and does not count as subscription-covered usage.
No. A subscription is for one person and is not a pooled team allowance. It cannot be shared, split, resold, or used to serve multiple users. Each person who wants subscription-covered usage needs their own account and subscription.
For shared or commercial team workloads, use pay-as-you-go access and the team controls instead of a personal subscription.
NanoGPT does not currently offer a separate pause action. To stop the next renewal while keeping access through the paid period, schedule a cancellation instead.
Open Subscription and choose Cancel subscription. Stripe subscriptions are managed through the Stripe Customer Portal; balance-funded subscriptions can be stopped from renewing directly on NanoGPT.
After cancellation is scheduled, access remains active until the displayed cancellation or current-period end date. If a Resume option is available before that date, use it to restore renewal. Canceling does not reset the weekly included-input quota.
Only the payment methods currently shown on the Balance or Subscription checkout are available. Availability can vary by region and payment method. Reopen the relevant page, confirm you are using the same NanoGPT session or account that made the payment, and check whether the card or crypto payment is still pending.
If a completed payment is not reflected, do not pay again immediately. Open a private support ticket from the affected session and include the payment method, approximate time, amount, receipt or transaction reference, and the Support Key shown by NanoGPT. Do not post full card details or private wallet credentials.
Yes. If your organization needs a written offer, quote, or pro-forma document for funding approval, grants, procurement, or internal purchasing, contact support with the billing entity, required reference numbers, requested credit amount, currency, and any deadline.
These documents are usually prepared for prepaid NanoGPT credits. Credits on NanoGPT are denominated in USD, so an offer can state the funding amount in another currency and clarify that the credited balance is converted at the exchange rate used at the time of payment or deposit.
Final VAT or tax details are stated on the invoice where applicable.
Yes. NanoGPT has a testimonials page with user reviews and links to original sources where available.
Yes. NanoGPT provides OpenAI-compatible API endpoints. Create an API key from the API page and use it with the documented base URL and model IDs.
NanoGPT works with many OpenAI-compatible clients. Create an API key, follow the dedicated integration guide when one exists, and use the exact base URL or endpoint required by that client.
The integrations documentation includes setup guides for JanitorAI, SillyTavern, OpenCode, OpenWebUI, Cursor, Cline, Codex CLI, Claude Code, and other tools. A third-party client can still have its own limitations, so first verify the API key, endpoint, model ID, subscription-versus-pay-as-you-go mode, and the client’s reasoning-field support.
Start with the HTTP status and machine-readable error code. Fix credentials for 401, add funds or disable paid extras for 402, check model permissions for 403, verify the model or resource ID for 404, and respect `Retry-After` for 429.
Timeouts and temporary server errors such as 408, 500, 503, or 504 can usually be retried with exponential backoff. A content-policy error is not a transient network failure and should not be retried unchanged. When contacting support, include the `X-Request-ID` response header when available.
Gibberish can look like random words or symbols, mixed languages, missing punctuation, a duplicated answer, repeated phrases, or an output loop that never reaches an answer. A reply may start normally and then degrade into nonsense. Stop the generation if it keeps running. This symptom can come from unsuitable generation settings, a very long conversation, or a temporary problem affecting the selected model.
First, retry the prompt in a new chat with little or no earlier context. If you use the API or a third-party client, choose “Model Default” or omit custom sampling settings such as `temperature`, `top_p`, `top_k`, `min_p`, `frequency_penalty`, and `presence_penalty`. A setting that works for one model may not suit another. If an image triggers the problem, confirm that the selected model supports vision and compare with a short text-only prompt. If thinking appears in the visible answer or the answer is duplicated, also check that the client supports the model’s reasoning format.
If the problem still occurs in a fresh chat with model-default settings, use Report response on that exact message when available, or open a support ticket. Include the model, approximate date and time with time zone, request ID when available, client or app, whether thinking or image input was enabled, and a short excerpt or screenshot. Never include an API key or other secret.
The Usage page shows account activity, costs, model usage, subscription usage, and media-generation records. Use its time range and model or API-key filters to narrow the results, and open a row for the available request or generation details.
Request/response body logging is currently available for new `/v1/chat/completions` requests made with an API key after logging is enabled for that key. The setting does not retroactively create logs and does not capture every NanoGPT API endpoint.
Usage and billing rows are an audit trail and cannot be selectively deleted from the Usage page. Stored request/response logs are separate. Disabling logging stops future captures and immediately deletes retained logs for that API key. The “Clear all account request logs” action deletes all stored account request logs without changing which API keys will log future requests.
Use the request’s output-token limit, such as `max_tokens`, `max_completion_tokens`, or the endpoint’s documented equivalent, to cap generated output. This is different from counting input tokens or changing the model’s context window.
For thinking models, reasoning tokens may consume part or all of that output budget. A very low limit can therefore produce a truncated answer or little to no visible answer. When debugging, omit the limit or leave enough headroom for both reasoning and the final response.
For supported reasoning models, use `reasoning_effort` or `reasoning.effort` with a documented level such as `none`, `low`, `medium`, or `high`. Not every model supports every level, so check the model details and API documentation.
Use the Support page to create a private ticket tied to your current session. The help assistant can also prefill a ticket draft, but it will only create the ticket after you confirm.
Use the Terms of Service and Privacy Policy for NanoGPT account, usage, payment, data-handling, and provider terms. Provider retention differs by route; Zero Data Retention applies only where the selected provider and route explicitly support it.
Do not assume every model or provider is ZDR. For a compliance or security-audit requirement that is not answered by the published policies, contact support with the specific control or evidence you need.
Use the newsletter signup on the Help page to get occasional NanoGPT updates, including major product changes and new model additions.
Newsletter emails are separate from account emails and include an unsubscribe link.
NanoGPT works with many model providers. If you offer models NanoGPT should add, or can offer better pricing or reliability, contact the team through support.
NanoGPT exists to make many AI models available through one pay-as-you-go interface. The goal is to reduce friction for people who want access to strong models without managing many subscriptions or provider accounts.
NanoGPT also supports privacy-conscious and crypto-friendly payment paths for users who cannot or do not want to use traditional subscription flows.
Get NanoGPT updates
Occasional emails for major improvements and new model additions. Unsubscribe anytime.