Silicon-based flow
A domestic AI infrastructure platform that offers multi-model inference, API calls, and computing power services.
Tags:AI training modelsWhat is silicon-based flow?
SiliconFlow is a company that provides infrastructure for running large models as well as enterprise-level MaaS platforms; its public cloud service for developers is called SiliconCloud. It brings together language, vision, image generation, speech, embedding, reordering, and video models from various manufacturers and open-source communities on a single platform, allowing users to access these capabilities by creating an API Key and making requests on a pay-as-you-go basis.
Unlike vendors that offer only a single proprietary model, SiliconCloud functions more like a cloud platform for multi-model inference. Developers can compare the prices of different models, as well as their capabilities in terms of context handling, tool integration, visual processing, inference performance, and speed, all within the same account and interface framework. They can then select the appropriate model for their needs, without having to purchase GPUs, deploy weights, or maintain inference clusters on their own.
Unified invocation of over 100 models
The current model hub displays over 100 popular models, including those from DeepSeek, Qwen, GLM, Kimi, MiniMax, Hunyuan, Step, BaiChuan, as well as projects from the open-source community. The categories covered are dialogue, coding, visual understanding, image generation and editing, speech recognition, speech synthesis, Embedding, Reranking, and video generation.
Each model page lists the model ID, context length, parameter size, input/output costs, and supported capabilities. The full model ID must be used when making a call;
The same model may include both regular connections and high-quality connections that start with “Pro”; the prices, concurrent usage capacity, and quality of service for these two types of connections can vary.
OpenAI-compatible API
SiliconCloud offers OpenAI-compatible interfaces; applications that use common OpenAI SDKs or allow customization of the Base URL can be migrated by simply replacing the interface address, API Key, and model ID. The official documentation also provides instructions on how to configure tools such as Kilo Code, Cherry Studio, Dify, LangChain, LlamaIndex, and Refly.
Compatibility does not mean that all parameters across different manufacturers are identical. The reasoning process, tool calls, multi-modal inputs, streaming outputs, and the maximum token limit vary depending on the model used.
Before going live, it is necessary to provide instructions regarding the reading capabilities of the target model, and to address issues such as unsupported parameters, as well as changes in the structure of the data returned by the model when it is taken offline.
Dialogue and reasoning models
The dialogue interface supports normal chatting, long texts, code, reasoning, stream-based output, as well as Function Calling for certain models. Reasoning models typically return the content of their reasoning in a separate reasoning_content field; the backend system should handle multiple rounds of context in accordance with the official guidelines, to avoid retransmitting lengthy reasoning histories verbatim.
There are significant differences in the capabilities of these models: some are suitable for low-cost classification and summarization, some focus on code handling and the use of Agent tools, while others support visual processing as well as very long contexts. It is not sufficient to choose a model based solely on its parameter size; it is necessary to compare factors such as accuracy, delay in generating the first token, overall latency, context capacity, maximum output length, and price as well.
Vision, images, and videos
Multimodal dialogue models can process inputs in the form of images or videos; image models enable text-to-image generation and image editing, while video models charge based on the output generated. The generation interfaces typically return file addresses that are valid for a short period of time – the official documentation states that such addresses remain valid for around 1 hour only, so applications must download the files promptly and store them in their own object storage systems.
Image and video generation is also influenced by factors such as size, duration, model used, concurrency levels, and content security policies. Asynchronous tasks should keep track of the request ID and the polling status, and handle cases of failure by providing refunds or allowing retries; temporary result addresses should not be used as permanent links to the content.
Voice, Embedding, and Rerank
The platform supports speech recognition, speech synthesis, and certain dialect models; it also offers various embedding and reranking models in both Chinese and English. Embedding is used to convert text into vectors, while reranking is used to reorder the search results – these are common components in knowledge bases and RAG systems.
The model center still offers several free embedding, reranking, OCR, translation, and speech models. The fact that they are free does not mean unlimited usage – there may be constraints related to concurrency, speed, waiting times, and the possibility that these services could become paid in the future.
The production system should have a paid backup model.
Current price pattern
| Package or version | Prices, quotas, and core benefits |
|---|---|
| GLM-5.2 | 8 yuan is inputted, with an output of 28 yuan per million tokens. |
| DeepSeek-V4-Flash | 1 yuan is entered, and 2 yuan per million tokens is output. |
| DeepSeek-V3.2 | 4 yuan is inputted, with an output of 6 yuan per million tokens. |
| Qwen3.5-35B-A3B | Within a 128K context, the cost is 0.40 yuan for input and 3.20 yuan per million tokens for output. |
| Image model | The common price is 0.10 to 0.30 yuan per piece, with some models available for free. |
| MOSS-TTSD and CosyVoice2 | The example price is 0.05 yuan per thousand UTF-8 characters. |
| Wan2.2 video | The price for examples of video generation from text and image generation into video is 2 yuan per example. |
SiliconCloud does not have a fixed monthly membership fee; instead, it relies on pre-funded balances along with usage-based charges. For dialogue models, the costs for input, output, and cache hits are calculated on a per million Tokens basis.
Images are charged per piece, audio is charged per thousand characters or according to the model’s rules, and videos are charged on a per-unit basis.
There is a large difference in prices among different models.
As of this verification, the example prices from the pricing center are as follows: GLM-5.2 costs 8 yuan for input and 28 yuan per million tokens for output; DeepSeek-V4-Flash costs 1 yuan for input and 2 yuan for output.
For DeepSeek-V3.2, the cost is 4 yuan for input and 6 yuan for output; for Qwen3.5-35B-A3B, it is 0.40 yuan for input when the volume is within 128K, and 3.20 yuan for output.
Long-context models may apply higher tiered pricing once the threshold is exceeded.
Currently, the cost for generating images with these models ranges from 0.10 to 0.30 yuan per image, with some models being available for free; for MOSS-TTSD and CosyVoice2, the cost is 0.05 yuan per thousand UTF-8 characters.
The examples for video generation from text and image in Wan2.2 cost 2 yuan each. These prices are subject to frequent changes; therefore, for official budgeting purposes, it is necessary to consult the current price information for the desired model, rather than relying on screenshots from articles.
Free models and complimentary balance
The platform’s pricing center marks some models as free, such as various Embedding, Rerank, OCR, translation, speech, and image models. These free models are suitable for testing, teaching, and low-cost RAG use, but their availability, speed, and scaling strategies may not meet the SLAs required for production environments.
New users, those who are invited, and participants in promotional campaigns may receive a bonus balance, but the official payment agreements do not specify a fixed amount that must be deposited upon registration. This bonus balance cannot be withdrawn, transferred, or converted into a receipt; its validity period and scope are determined by the rules of the respective campaign. Therefore, the practice of offering a fixed amount as a bonus for registration should not be considered a permanent entitlement.
Top-up, balance, and refunds
Before making an online top-up, individuals must complete identity verification. The only online payment method listed in the top-up agreement is WeChat Pay. The amount deposited through this payment becomes part of the account balance, and this balance has no expiration date; however, it cannot be transferred or given as a gift.
The amount of the gift is managed separately.
Within 360 days after the online payment is completed, any unused balance can be refunded automatically following approval; a handling fee may apply. The amounts that have already been spent as well as any complimentary balance cannot be refunded.
In principle, only one refund is allowed per online top-up; for transfers between businesses, it is necessary to contact support for processing.
Bills and invoices
The console provides expense bills, allowing one to view consumption by model and by number of calls. Invoices can only be issued for amounts that have actually been spent; the balance remaining after a top-up but not yet used cannot be invoiced.
A complimentary balance cannot be issued as an invoice.
With personal verification, only ordinary personal invoices can be issued; with corporate verification, invoices on behalf of the company can be requested in accordance with the relevant rules.
Production projects should have their usage quantified based on API Keys, environment, or business units, and alerts for low balances should be set in place. Testing, development, and production environments should not share the same key, as this makes it difficult to track abnormal calls and allocate costs appropriately.
API Key security
An API Key grants permissions for account usage; it must be stored in server-side environment variables or a key management system, and must not be included in website front-ends, mobile application install packages, screenshots, or public code repositories. If a leak is detected, it should be deleted immediately and recreated.
It is recommended to create separate keys for each application, to set up monitoring of costs and request logs, and to avoid recording the users’ original sensitive input statements. If the client must make calls over the public network, it should first access its own backend, which will then forward the requests to SiliconCloud.
Throttling and concurrency
The platform sets limits on RPM, TPM, RPH, RPD, or concurrent usage based on the model, connection type, account status, and real-time load; official announcements also adjust these policies according to the available resources. Free models as well as popular new models are more likely to experience waiting times or have their usage limited.
To handle 429 errors, timeouts, and service errors, mechanisms such as exponential backoff, circuit breaking, queues, and backup systems should be employed; it cannot be assumed that a request will always succeed. For high-concurrency production needs, the Pro plan can be chosen, or contact enterprise services to arrange customized concurrency settings and SLAs.
Model fine-tuning
The SiliconCloud console provides a tool for model fine-tuning; separate charges are applied for training and inference after fine-tuning. Once a base model is selected, the page will display the corresponding costs for training and inference.
Fine-tuning is suitable for fixing formats, industry-specific expressions, and optimizing task behaviors, but it is not appropriate for injecting frequently changing knowledge in real time.
Before uploading training data, personal information, client secrets, and any unauthorized content should be removed, while a validation set should be retained to check for overfitting. In scenarios where knowledge needs to be updated frequently, RAG and Embedding methods are generally the preferred choices.
Enterprise-grade MaaS and privatization
Enterprise-grade MaaS covers the management of heterogeneous computing resources, model training, inference deployment, fine-tuning, monitoring, and resource recycling; it is compatible with various architectures as well as domestic chips. Its intended applications include scenarios in areas such as AI computing centers, energy, manufacturing, transportation, and telecommunications operators.
For enterprises that require sensitive data to remain within the local network, already possess GPUs or domestic computing resources, and need specialized capabilities as well as compliance controls, they can consult about localized deployment options. Enterprise solutions come with customized pricing, and the contract should specify details regarding hardware compatibility, model licensing, upgrades, SLAs, security assessments, and the scope of maintenance services.
OneDiff and inference acceleration
OneDiff is an open-source acceleration library for diffusion models provided by Silicon Base Flow, designed to optimize the inference process for generation models such as Stable Diffusion. It operates at a different level compared to the SiliconCloud API in the cloud: the former is an acceleration tool that developers can deploy, while the latter offers hosted multi-model services.
When using OneDiff, it is still necessary to provide one’s own GPU, model weights, and operating environment, in addition to complying with the licenses associated with each model. It is not possible to copy all of SiliconCloud’s commercial models, scheduling systems, and enterprise consoles to a local setup.
BizyAir and ComfyUI
BizyAir provides ComfyUI nodes that enable workflows to take advantage of cloud-based generation capabilities even in environments without a local GPU. It is suitable for integrating SiliconCloud models into node-based image processing workflows and combining them with other ComfyUI nodes.
Cloud nodes consume the platform’s balance, and the data generated is processed through cloud services. When dealing with unpublished materials, portraits of individuals, or customer files, it is necessary to review the data processing policies; one should not assume that all operations are carried out locally just because ComfyUI itself is open-source.
GitHub and the open-source status
The official GitHub organization for Silicon-based Flow has made available nodes related to OneDiff, BizyAir, siliconcloud-cookbook, and ComfyUI, as well as several engineering projects; it also maintains forks of some upstream projects. The licensing terms for each repository vary, so it is necessary to check them individually.
The SiliconCloud cloud platform, as well as the services related to model routing, billing, scheduling, and comprehensive inference, are not fully open source. Therefore, the accurate description should be “the cloud platform is not open source; however, several open source projects are provided by the vendor”, and it cannot be claimed that the entire SiliconCloud system can be used privately at no cost.
Tutorial for silicon-based flow usage
Complete a basic task.
- Register for silicon-based streaming and create an API Key intended solely for use in a testing environment;
- Select a model based on input type, context, quality, speed, and price;
- First, invoke over 100 models in a unified manner to complete the minimal request and check the returned structure;
- Then, use dialogue and reasoning models to test streaming output, parameters, and abnormal response handling.
- Record Tokens, number of calls, latency, error rate, and cost per call;
- Move the key to the server-side key manager before integrating it into the actual application;
Create reusable professional workflows
- Different keys and quotas are used for development, testing, and production environments;
- Unified invocation via over 100 models, as well as representative evaluation sets for dialogue and reasoning models alongside vision, images, and videos;
- Set timeout, concurrency, retry, throttling, and budget limits;
- Perform checks on the output regarding facts, security, format, and sensitive information;
- Monitor changes in model version, price, latency, and failure rate;
- Prepare plans for downgrading the model, implementing circuit breaking, and taking manual control;
Which users is it suitable for?
- Developers who wish to quickly compare and switch between domestic and open-source large models using an API;
- Teams for building chat, code, Agent, RAG, image, voice, or video applications;
- Startups and individual developers who do not want to set up their own GPU inference clusters;
- Knowledge base projects that require embedding, reranking, and multimodal capabilities;
- Companies that require domestic computing power for adaptation, localized MaaS solutions, and dedicated services.
Product advantages
- Unified accounts and APIs support over 100 models;
- It is compatible with the OpenAI ecosystem, and the process of migration as well as the configuration of third-party tools are relatively straightforward.
- It covers text, visuals, images, audio, vectors, reordering, and video at the same time;
- Pay-as-you-go pricing with multiple free models available;
- Offers model fine-tuning, Pro plans, and enterprise-specific private MaaS solutions;
- The team is responsible for maintaining open-source projects such as OneDiff, BizyAir, and Cookbook.
Restrictions and Precautions
- The list of models, their prices, versions, and usage limits change frequently; older models may be discontinued or seamlessly updated to new versions.
- Applications that rely heavily on output behavior should fix the model ID, establish regression tests, and verify them after any updates are released.
- The aggregation API does not eliminate the model’s inherent issues related to illusions, copyright, as well as security and compliance risks.
- Free lines do not guarantee long-term stability;
- The temporary address of the generated file needs to be saved promptly.
- A Key leak can result in a loss of balance;
- Medical, legal, financial, and government-related applications must incorporate additional manual review processes as well as sector-specific security controls.
Frequently Asked Questions
Is silicon-based streaming free?
The platform operates on a pay-as-you-go basis, but the Model Center offers several free models, and promotions may also provide additional credits. The list of free models and the amounts of credit available can change; it does not mean that all models are free.
How much is the silicon-based flow API?
There is no fixed unit price. The cost of conversations is calculated based on input, output, and cached tokens, with current rates ranging from a few cents per million tokens to several dozen yuan.
Images typically cost starting from 0.10 yuan, while some videos cost around 2 yuan each; it is necessary to check the real-time price list for models.
Is the silicon-based solution compatible with the OpenAI API?
It is compatible with common OpenAI interfaces and SDK configurations, but inference, multi-modal capabilities, tool calls, and model-specific parameters still require adaptation in accordance with the official documentation.
Does the top-up balance expire?
The actual balance paid for top-ups has no expiration date; the complimentary balance is governed by the rules of the campaign.
A refund can be requested for the unused balance from online payments within 360 days after top-up, subject to approval.
Is the silicon-based flow open source?
The core cloud platform of SiliconCloud is not open source, but the company does make projects such as OneDiff, BizyAir, and Cookbook available under an open source license. An open source repository does not equate to a complete cloud platform that can be hosted on one’s own.
Guigong Network Security Registration No. 45132202000164