Skip to main content
Reach out to our support team with details about your use case, expected volume, and any latency or throughput requirements. We’ll review your needs and can raise your limits accordingly.
Implement exponential backoff in your client code. When you receive a rate limit error (429) or a server error (503), wait before retrying.

How are tokens counted?

Tokens are pieces of text that our models process. A token is roughly 4 characters for English text. Both input and output tokens count toward your usage limits.
Yes. The full documentation index is published at https://docs.inceptionlabs.ai/llms.txt. Point your LLM, IDE assistant, or agent at that URL to discover every page in the docs. The docs site does not block crawlers or agent traffic.
No. Mercury 2 and Mercury Edit 2 accept text input only. Image generation and image input are not supported. See /get-started/models for supported input formats per model.