> ## Documentation Index
> Fetch the complete documentation index at: https://docs.llumo.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# LLUMO AI - Cut LLM Cost by 50%

> LLUMO compresses your tokens to build production ready AI at 50% cost and 10x speed.

<img className="block dark:hidden" src="https://mintcdn.com/llumoai/KG4A3Af7fc7C2-tf/images/hero-light.svg?fit=max&auto=format&n=KG4A3Af7fc7C2-tf&q=85&s=ede0f864eaf3fa4bee55375f5d1fc610" alt="Hero Light" width="700" height="320" data-path="images/hero-light.svg" />

<img className="hidden dark:block" src="https://mintcdn.com/llumoai/KG4A3Af7fc7C2-tf/images/hero-dark.svg?fit=max&auto=format&n=KG4A3Af7fc7C2-tf&q=85&s=e18b1156db66a0b3a15871b2b2e33e9d" alt="Hero Dark" width="700" height="320" data-path="images/hero-dark.svg" />

[LLUMO AI](https://app.llumo.ai) is a plug and play API tool which helps you reduce LLM inference cost by more than 50%
and speed up inference by 10x. It is a simple and easy to use API that can be integrated into your backend code just
before you send calls to LLM.

[LLUMO AI](https://app.llumo.ai) helps you and your team

* Compressed prompt & output tokens, to cut your AI cost with augmented production level AI quality output.
* Efficient chat memory management slashes inference costs and accelerates speed by 10x on recurring queries.
* Monitor your AI performance and cost in real-time to continuously optimize your AI product.
