AI21
AI21: an intelligent tool focused on AI programming.
Tags:AI programming toolsWhat is AI21 Maestro?
AI21 Maestro is a dynamic planning, orchestration, and optimization system developed by AI21 Labs for production-grade AI Agents. It selects the appropriate models, tools, and execution paths at runtime based on tasks, requirements, and budget constraints, and it verifies and adjusts the intermediate results as well as the final output.
A one-sentence summary
Maestro can combine corporate documents, web pages, models, and tools to create trackable knowledge agents, which automatically seek optimal solutions in terms of quality, cost, and response time.
Product positioning
- Designed for high-value, data-intensive enterprise tasks.
- Used to create and optimize RAG knowledge Agents.
- It emphasizes dynamic programming at runtime rather than fixed paths.
- Improve controllability by requiring verification and execution tracking.
- It supports unified orchestration of AI21 and third-party models.
- There are existing APIs and SDKs, as well as the possibility of configuration through the platform.
Core functions
- Break down complex tasks into multi-step execution plans.
- Select the appropriate model, as well as search and external tools.
- Explore multiple candidate paths in parallel.
- It is calculated based on quality, cost, and delay budget allocations.
- Check the fulfillment of requirements step by step.
- Automatically retry or correct after detecting an issue.
- Execution graphs, confidence levels, and verification reports are provided.
- Save the Agent configuration for subsequent API calls.
Dynamic programming and orchestration
Each time it runs, Maestro constructs an execution tree that consists of model calls and tool operations, rather than always using a fixed sequence of steps. The plan changes in response to requirements, data, and budget constraints.
- Analyze the task objectives and explicit requirements.
- Select the retrieval and reasoning steps that need to be carried out.
- Compare the expected quality and cost of different paths.
- Adjust the next step during operation.
- Stop when the additional calculated value is insufficient.
- The complete execution trajectory is output for review.
Multi-path competition
- Generate multiple candidate strategies for complex problems.
- Parallel paths are allowed to use different models or tools.
- Compare the degree to which each candidate result meets the requirements.
- Eliminate low-quality or high-cost approaches.
- Keep the best results.
- Parallel computing increases actual costs and delays.
Built-in verification and automatic correction
Maestro will score the output based on the user-defined requirements regarding content, style, format, and constraints. If the result does not meet the desired standards, further searches, reasoning, or adjustments can be made until the requirements are satisfied or the budget is exhausted.
- A separate score is returned for each requirement.
- Provides an overall satisfaction score.
- An error was found in the intermediate step.
- Prevent errors from being passed on to subsequent processes.
- Automatically attempts to fix failed results.
- The verification results still need to be verified through manual sampling.
Budget control
| Budget level | Execution characteristics | Suitable for tasks | Main choices to make |
|---|---|---|---|
| Low | Single-step, minimum computation | Simple or time-sensitive tasks | Fast speed, low cost, and minimal verification required. |
| Medium | Multi-step process with moderate input requirements | Routine corporate knowledge tasks | Balance between quality, speed, and cost |
| High | Multiple strategies and multiple rounds of verification | Complex or high-risk tasks | Higher reliability, but increased costs and delays |
Optimization of quality, cost, and delays
- Set an acceptable quality threshold for the task.
- Limits the computing budget that can be used per execution.
- Set delay targets based on the scenario.
- Observe the quality cost curves for different configurations.
- Route simple tasks to the lighter execution path.
- Perform calculations only when the expected quality improvement is worthwhile.
- atribute expenses by team, code repository, and Agent.
RAG knowledge retrieval
Maestro is built on AI21’s RAG Engine; it enables the combination of file searching, web searching, and subsequent question handling at the planning level. Answers can be grounded in corporate documents, but RAG still cannot completely eliminate hallucinations.
- Files are automatically parsed and indexed after being uploaded.
- Retrieve relevant content from the specified document.
- Supplement public information through web searches.
- Continuous follow-up questions and clarifications are supported.
- Use the search results for multi-step reasoning.
- Suitable for internal knowledge and customer applications.
Supported file types
- PDF document.
- Microsoft Word document.
- Plain text file.
- Markdown file.
- HTML file.
- Cloud data source connections must be confirmed based on the account’s capabilities.
- Scans, complex forms, and images should be tested separately for their parsing quality.
Structured RAG and data extraction
- Extract structured information from complex documents.
- Identify table, field, and document relationships.
- The extraction results will be used in subsequent Agent steps.
- Suitable for contracts, reports, and technical documents.
- It can be used for data migration and standardization.
- High-risk fields must be verified against the original text.
Model-agnostic orchestration
Maestro can use AI21’s own models, as well as third-party models managed by AI21 or for which keys are provided by the users themselves. Teams can specify the models to be used, or they can let the system choose them automatically based on the task at hand.
- Supports AI21 Jamba series models.
- Some OpenAI models can be used.
- Some of Anthropic’s models can be used.
- Some Google models can be used.
- Reduce differences in upper-layer integration through a unified interface.
- Third-party available models will be updated as the service evolves.
- The use of the model is still subject to the terms set by various suppliers.
BYOK comes with its own keys
- Configure third-party model credentials in the platform.
- The current document lists OpenAI, Anthropic, and Google.
- Reference the configured model ID when the Agent is running.
- Third-party fees are calculated by the respective suppliers.
- Special projects and restricted keys should be used.
- Regularly rotate and monitor abnormal calls.
- Do not write the key directly into the client code.
Remote MCP tool
Maestro can act as an MCP client to connect to remote MCP servers that are accessible over the Internet, thereby enabling the use of external tools, services, and data. This documentation covers remote servers only; it does not address local MCP servers that operate solely on a single device.
- It enables connection to remote MCP services that are accessible over the Internet.
- Use common authentication methods to protect the calls.
- Add external tools to the Agent execution plan.
- It supports workflows that require multiple steps.
- The tool’s return value can be used for subsequent verification.
- External write operations should require manual approval.
Save and reuse Agents
- Save the Agent name and description.
- Save system instructions and business boundaries.
- Save tool and model configurations.
- Reference it directly in subsequent API calls.
- Reduce the repeated transmission of large configurations.
- It helps to standardize the team’s workflow.
- Version and regression testing should be conducted when modifying the configuration.
Execution tracking and observability
- View the visual execution diagram.
- Understand the models, tools, and the sequence of retrieval calls.
- Check the validation scores for each requirement.
- Steps with failed positioning or low quality.
- Analyze the sources of cost and delay.
- Compare different budget configurations.
- Retain necessary operational records for auditing purposes.
Supported response languages
The API currently supports output in Arabic, Dutch, English, French, German, Hebrew, Italian, Portuguese, and Spanish; English is used by default. Chinese is not included in the current list of supported languages.
Which users are it suitable for
- Form an AI engineering team to develop corporate knowledge agents.
- Financial and legal institutions that require high-accuracy RAG.
- Companies that handle a large number of contracts, reports, and technical documents.
- Developers who wish to unify multiple model suppliers.
- Platform teams that need to control Agent costs and latency.
- Regulated organizations that wish to retain records of execution and verification are encouraged to do so.
- Companies that put AI research and analysis processes into production.
Typical use cases
- Generate financial research reports.
- Automatically prepare RFP responses.
- Conduct due diligence on mergers and acquisitions.
- Review the contract portfolio and key terms.
- Analyze the investment prospectus and financial documents.
- Summarize the results of clinical trials.
- Identify technical issues with high-value equipment.
- Extract the claim elements from the patent.
- Convert data from the old system into a standard structure.
It’s not very suitable for which situations
- Just a person for a simple back-and-forth conversation.
- Scenarios where the lowest latency is sought and no verification is required.
- Organizations that must operate entirely offline and locally.
- Applications that require enumerations to be output in native Chinese are needed.
- Teams that do not have permissions for engineering capability management tools.
- Developers who wish to obtain the complete Maestro source code.
- Projects with pay-per-use billing and fluctuating costs are not acceptable.
How to use the API
- Create a Maestro run to execute the task.
- Pass in user input and explicit requirements.
- Search for configuration files or web resources.
- Specify a model or use automatic selection.
- Set a budget of low, medium, or high.
- Select the tracking information that needs to be returned.
- Read the output, the required score, and the overall rating.
Price and trial quota
The AI21 documentation states that the platform charges primarily based on the number of tokens or requests to specific endpoints. New accounts are granted a 10-dollar trial credit valid for three months, which can be used for APIs, SDKs, and the Playground. Once this credit is used up or expires, it is necessary to provide a valid payment method.
| Cost type | Public rules | Precautions |
|---|---|---|
| Trial for new accounts | 10 dollars, valid for 3 months | It’s not a long-term free plan. |
| AI21 platform usage | Charging by token or endpoint | The input and output prices may differ. |
| Maestro is running. | Affected by budget, steps, models, and tools | A high budget usually increases costs. |
| Third-party hosting model | Charged based on AI21 or the corresponding configuration. | It is necessary to check the current price list for the model. |
| BYOK model | Third-party suppliers charge additional fees. | The operating costs of AI21 still need to be determined. |
| Cloud platform deployment | Priced by platforms such as AWS or Azure | Charging may be based on tokens, instances, or time. |
| Large-scale enterprises | Customized solutions | Contact AI21 for a business quote. |
Factors affecting costs
- Number of input and output tokens.
- The selected model and supplier.
- Budget level: Low, Medium, or High.
- Execution path and number of verification loops.
- Web search, file retrieval, and invocation of external tools.
- Number of retry attempts on failure and automatic correction times.
- Fees for third-party cloud and data services.
How to control costs
- Simple tasks start with a Low budget by default.
- Reduce the scope of file searches and website domain names.
- Tasks will be explicitly routed to lighter models.
- Enable strong verification only for critical requirements.
- Monitor the usage field in the API response.
- Set budgets by team, application, and Agent.
- Use the sample set to compare quality and unit cost.
Tutorial on Creating a RAG Agent
- Create an AI21 account and securely save the API key.
- Upload a small number of authorized test documents.
- Wait for the file to be automatically parsed and indexed.
- Define the agent’s identity, tasks, and prohibited areas.
- Add file search and specify files or tags.
- Representative issues arising from executing projects with a low budget.
- Check citations, responses, and verification reports.
- Save the Agent and integrate it with the testing application using the API.
Requirements configuration tutorial
- Write the task objectives and evaluation requirements separately.
- Use clear names for each requirement.
- Describe the content, format, tone, and guidelines.
- Avoid conflicting or unverifiable requirements.
- Set an appropriate budget and carry out the tests.
- View the score for each requirement and the reasons for failure.
- Re-run the regression tests after modifying the requirements or tools.
MCP Tool Integration Tutorial
- Verify that the MCP server can be accessed securely over the Internet.
- Review the capabilities of the tool, its parameters, and the risks associated with writing data to it.
- Configure dedicated identities and least-privilege authentication.
- Add a remote service to the Maestro tool parameters.
- Use read-only and test data to verify the calls.
- Add manual approval for sensitive actions.
- Monitor call logs, as well as costs related to errors and exceptions.
Evaluation and deployment process
- Create datasets of real tasks and standard answers.
- Define accuracy, requirement satisfaction rate, and latency metrics.
- Test different models and budget levels separately.
- Failed checks, rejected responses, and incorrect tool calls.
- Calculate the total cost for each qualifying result.
- Launch with a low volume of traffic and retain manual review.
- Continuously monitor data, model, and task drift.
Security and data governance
- Only authorized corporate information should be uploaded.
- Assign permissions to files, Agents, and projects.
- BYOK uses independent and revocable keys.
- Apply the principle of least privilege to remote tools.
- Manual confirmation is required for sensitive write operations.
- Define the retention periods for operation logs and documentation.
- Verify the trust center and the scope of the contract before making a purchase.
- Delete test data and regularly rotate credentials.
Agent reliability risk
- The verification model itself can also make mistakes.
- A high score does not necessarily mean that the information is completely accurate.
- Omissions in the search will limit the final answer.
- Parsing complex documents can damage tables and context.
- Third-party models and tools introduce additional faults.
- Dynamic execution leads to fluctuations in costs and delays.
- High-risk tasks still require review by domain experts.
Product advantages
- The execution path is dynamically selected at runtime.
- Put quality, cost, and delay under the same control framework.
- Score and revise each specific requirement.
- Supports files, web pages, and remote MCP tools.
- AI21 and third-party models can be orchestrated in a unified manner.
- BYOK is supported to reduce model locking.
- Execution graphs and structured verification reports are provided.
- The official Python and TypeScript SDKs facilitate integration.
Product restrictions
- The full Maestro platform is not open-source software.
- High budgets and multiple paths increase costs and delays.
- Verification cannot eliminate all factual and reasoning errors.
- The current list of supported response languages does not include Chinese.
- Local MCP servers are not within the current scope of support.
- Complex permissions and tool calls require engineering governance.
- The price is influenced by various models and services.
- The public page does not offer any simple, fixed monthly subscription plans.
- Sensitive corporate data requires additional compliance assessments.
GitHub and the open-source status
AI21 Labs has an official GitHub organization that makes the Python and TypeScript SDKs available publicly; the relevant clients are licensed under the Apache 2.0 license and include examples of how to use Maestro. The fact that the SDKs are open source does not mean that the source code for Maestro’s dynamic programming, validation engine, and hosting platform is also made available.
Basic information
| Project | Content |
|---|---|
| Tool name | AI21 Maestro |
| Development company | AI21 Labs |
| Tool type | Enterprise AI Agent planning, RAG, and optimization systems |
| Budget level | Low, Medium, High |
| Model | AI21, Managed Third Parties, and BYOK |
| Trial | 10 dollars, valid for 3 months |
| APIs and SDKs | REST, Python, and TypeScript |
| MCP | Supports remote servers |
| Open-source status | The platform is not open-source; the SDK is licensed under Apache 2.0. |
Recommendation score
Recommendation score: 4.6 / 5. AI21 Maestro is suitable for enterprise AI teams that need to deploy high-accuracy RAG solutions and handle complex knowledge-related tasks in a production environment, while also keeping costs and latency under control.
The final selection should be made by verifying the accuracy, the quality of interpretation, the reliability of the validations, the requirements related to Chinese language, as well as the actual cost of each viable option, using one’s own dataset.
Frequently Asked Questions
What is AI21 Maestro?
It is a dynamic planning system used for creating, orchestrating, and optimizing production-grade RAGs and knowledge agents.
Is Maestro free?
New accounts come with a $10 trial credit valid for three months; after that, payment is required based on usage.
Is third-party modeling supported?
It supports third-party models managed by AI21, as well as BYOK configurations from OpenAI, Anthropic, and Google.
Is Chinese supported?
The response_language enum available in the current API does not include Chinese; other output methods should be tested using actual tasks.
Is MCP supported?
Internet-accessible remote MCP servers are supported; currently, the documentation does not cover MCP servers that operate only locally.
What files can be uploaded?
It supports PDF, DOCX, TXT, Markdown, and HTML.
Is a higher budget necessarily more accurate?
It usually increases the number of paths and verification attempts, but it does not guarantee that all tasks will be more accurate.
Is Maestro open source?
The platform is not open-source; the official Python and TypeScript SDKs are licensed under the Apache 2.0 license.
Is it suitable for personal conversations?
It is more suitable for developers and corporate knowledge workflows; it may be too complex for ordinary conversations.
Guigong Network Security Registration No. 45132202000164