less to LLMSend context.
Aether Compress optimizes your context before the model call. It reduces unnecessary input token load, preserves usable information, and ensures your existing AI workflow runs more efficiently.
working before modela context optimization layer.
In an LLM call, the model not only asks the question, but also It processes all the context you send. this context As it grows, the amount of input tokens also increases.
Aether Compress is placed between your application and the model. It first optimizes the context and produces a smaller context. The next step is entirely under your control.
Part of the model costIt comes from the payload you send to the model.
Greater context always means more useful information. It doesn't come. Repetitions, auxiliary content and low relevance to the query The sections are also transferred to the model as input tokens.
As the input context grows, the amount of tokens sent and charged to the provider increases.
Unnecessary content consumes context space that truly necessary information could use.
What may seem like a small difference in a single request turns into large total token amounts at high call volume.
Context enters.Optimized context appears.
Aether Compress does not try to generate a model response. Its task is to make the input context more efficient before reaching the LLM.
You send the existing context in your application to the Aether API.
If there is no query, General optimization is applied, and if there is a query, Query Oriented optimization is applied.
Aether identifies the load that can be reduced while preserving the available information.
Optimized context returns directly to your application.
Query or not.
There are two different ways to use context. You can use optimization format.
General contextOptimize and protect.
context without a specific user question General mode when you want to optimize you can use
Context'iFocus on the query.
If the user's question is known, the query can also be added to the optimization request. In this way, Aether processes the context taking into account the information needs of that query.
Another modeleklemiyoruz.
Making a new LLM call to reduce context cost may mean moving the cost elsewhere. Aether Compress does not base its optimization process on a separate LLM call.
Aether Compress does not send a request to another AI model to optimize the context.
The optimized context can be used in the next stage with the model or provider you prefer.
You don't have to change your model, prompt architecture, or the main LLM layer of your application.
You can measure the result by comparing the amount of original and optimized tokens in each request.
Just get the context.Continue until the answer if you want.
Aether Compress is the optimization layer of the product. Aether Generate is the optional usage method that combines this layer with your AI provider in a single call.
Optimized context returns directly to your application. You decide which model, when and how to use it. You determine.
The context is first optimized, then its own provider is sent to the model through your account and the final answer rotates in a single flow.
With a single API calladd it to your stream.
It is sufficient to add Aether Compress between your current context creation stage and your LLM call.
Review documentationconst response = await fetch(
"https://api.aether.tr/v1/compress",
{
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Idempotency-Key": crypto.randomUUID(),
"Content-Type": "application/json"
},
body: JSON.stringify({
context,
query
})
}
);
const result = await response.json();
console.log(result.context);If the context is growing,It could be an area of optimization.
Aether focuses on the context structure sent to the LLM rather than a specific industry.
Optimize for long conversation histories and repetitive context before model calling.
Make the document fragments that come as a result of retrieval more efficient according to the user query.
Reduce the burden of tool outputs, task history, and accumulated agent context.
Do not carry long reports, contracts and corporate documents in their entirety with every request.
Optimize conversation history, customer information and help center context.
Minimize large contexts containing source code, logs, error output and technical documentation.
Not just smaller.Information must also be protected.
When evaluating Aether Compress, we do not only measure token reduction. We also separately verify that the optimized context retains the necessary information.
Check out benchmark resultsIf there is no safe reductionDoesn't have to reduce.
The goal is not to produce a smaller output under all circumstances. More important than reducing the context that needs to be preserved, protection of usable information.
Not the marketing average,Look at your own context.
Each application has a different context structure. The most accurate way to see the real value of Aether is to measure it on your own production inputs.
If the context cost actually exists.
If your requests carry conversation history, documents or large system context.
In systems where the same optimization is repeated across thousands or millions of requests.
If the input is already at the level of a few tokens, the space that can be reduced is naturally limited.
If your context has been aggressively minimized before, the additional gain may be lower.
Submit your context.See the difference in your own data.
Test Aether Compress on real inputs that resemble your production context with 1 million free tokens.
