AETHER · Frequently Asked Questions

If you have questionsThe answer may be here.

Aether Compress brings together in one place the questions we encounter most frequently about API usage, Aether Generate, pricing, verification, and security.

Genel
What exactly is Aether Compress?
Aether Compress is a context optimization layer that operates before the LLM call. It processes the context your application will send to the model, reduces unnecessary input load that can be minimized, and returns the optimized context.
Genel
Is Aether an artificial intelligence model?
No. Aether Compress is not an LLM and does not call another AI model to optimize the context. It is a separate optimization layer running in front of your existing AI architecture.
Genel
Which models can Aether be used with?
Aether Compress generates optimized context. Therefore, the model or provider you use in the next stage is independent of Aether Compress. You can use the optimized context in your existing AI workflow.
Genel
Do I need to change my current AI architecture?
You do not need to change your basic model architecture. Aether Compress is placed between the stage where the context is created and the LLM call. Your application receives the optimized context and continues with its current flow.
Compress
Does Aether Compress make an LLM call while running?
No. When you only use Aether Compress, no request is sent to any AI provider. The context is optimized and the result is returned directly to your application.
Compress
Is it mandatory to send a query?
No. If you do not send a query, General optimization is used. If you send query context, query will be considered It is optimized in a focused manner.
Compress
What is the difference between General and Query-Aware?
General mode is a context that is not tied to a specific question. focuses on maintaining its overall availability. Query-Aware In this mode, the user query is also included in the optimization input. and the context is taken into account by taking into account the information needs of that query. is processed.
Compress
Does Aether reduce the context with every request?
No. The goal is not to produce smaller output under all circumstances. context when there is no space that can be safely reduced It may return unchanged or with very limited change.
Compress
Is 40.65% token reduction guaranteed for every request?
No. 40.65% is the average input reduction measured on a validation set of 100,000 scenarios for Aether Compress v1.0. The actual result may vary depending on the structure, length of the context, and optimization mode.
Compress
What is the highest measured reduction?
Highest input measured in a validation set of 100,000 scenarios The reduction is 52.67%. This value is the expected or It is not a guaranteed rate.
Generate
What is Aether Generate?
Aether Generate first optimizes the context with Aether Compress, then sends it to the AI provider you select and returns the final model response to your application in a single stream.
Generate
What is the difference between Compress and Generate?
Compress It only does context optimization and optimized Returns context. Generate using your own AI provider after this optimization It also produces the final answer.
Generate
Do I need to provide my own LLM API key for Aether Generate?
Yes. BYOK in Generate usage, i.e. Bring Your Own Key model is used. Your supported provider account and API You provide your key.
Generate
Is the AI provider's usage fee included in the Aether plan?
No. The Aether plan covers the usage of input tokens processed through Aether. The model usage cost of an external AI provider will be billed separately to your own provider account.
API
How do I integrate Aether Compress?
You send the context to the Aether API before your existing LLM call. The API returns the optimized context, which you can then use in your existing model call.
API
Is the SDK required to use Aether Compress?
No. Basic integration can be done via HTTP API. The application language or framework you will use depends on the API. integration as long as it can send a standard HTTP request can be realized.
API
Is there a limit to the amount of tokens I can send in a single request?
Yes. The single request token limit varies depending on the plan you use. 25K in Free plan, 50K in Developer, 100K in Pro, 200K input token in Studio and 500K in Business There is a limit.
API
Does the rate limit change according to plans?
Yes. Minute request and concurrent request capacity plan increases as it rises. Detailed limits on the Pricing page. you can see.
API
Should I use Aether instead of the RAG system?
No. It does not replace the Aether retrieval system. Your RAG system fetches the relevant content, and Aether Compress is placed between this retrieval result and the LLM call to optimize the context.
Pricing
Is there a free plan?
Yes. The free plan includes 1 million input tokens per month. to get started no credit card required.
Pricing
How is the token quota calculated?
The plan quota is calculated based on the number of input tokens processed through Aether Compress.
Pricing
Which plans include Aether Generate?
Aether Generate is available in paid plans. The Free plan only includes the use of Aether Compress.
Pricing
Which model should I use with very high token volume?
For monthly usage of 5 billion tokens or more in your own company, a Technology License can be considered. If you want to offer Aether to your own customers, the Distribution Partner model is a more suitable structure.
Pricing
What is the difference between Distribution Partner and Technology License?
A Technology License is for using Aether in your own high-volume operations. A Distribution Partner is aimed at companies that want to offer Aether to their customers and manage a sub-customer structure.
verification
How was Aether Compress tested?
Aether Compress v1.0 was tested with a local validation set of 100,000 scenarios. In these tests, information preservation and input reduction were evaluated together.
verification
What does 100% accuracy mean?
Across all 500 scenarios selected for external model validation Maintenance of the required response was measured. This result tested belongs to scenarios and is universal for all possible inputs without limit. does not mean warranty.
verification
What does 3.48 ms mean?
3.48 ms is the average Aether Compress processing time measured in the verification environment. Actual latency may vary depending on hardware, infrastructure, request size, and operational environment.
verification
Where can I see benchmark results?
Detailed measurement results, methodology and measurement limits Benchmark is published on the page.
Security
Where can I review Aether's security approach?
Our security approach includes access control, infrastructure security and Details about the security notification process Security You can review it on the page.
Security
Where are the policies regarding personal data?
For general privacy approach Privacy Policy, Regarding personal data processing processes in Türkiye for information KVKK Information Text incelenebilir.
Still Have Questions?

Answer togetherLet's find it.

You can contact us about integration, high-volume usage, licensing, or the suitability of Aether for your architecture.