Nano Banana Pro
Nano Banana Pro: a smart tool focused on AI-powered image processing.
Tags:Common AI image toolsWhat is Nano Banana Pro?
Nano Banana Pro is a professional-grade image generation and editing model developed by Google DeepMind; its official name is Gemini 3 Pro Image. It leverages the reasoning capabilities, world knowledge, and multimodal understanding of Gemini 3 Pro for complex visual design tasks.
This model focuses on enhancing the consistency of text, infographics, product prototypes, multiple reference images, and the brand identity within images, and it supports output in up to 4K resolution. It remains the option for handling complex tasks within the Nano Banana family, but it is not the only image processing model available nor the fastest one.
Name and model code of Nano Banana
| Common names | Formal model | API model code | Positioning |
|---|---|---|---|
| Nano Banana 2 Lite | Gemini 3.1 Flash Lite Image | gemini-3.1-flash-lite-image | Fastest speed, lowest cost, and high throughput |
| Nano Banana 2 | Gemini 3.1 Flash Image | gemini-3.1-flash-image | Speed, 4K, text, and the core tools for editing |
| Nano Banana Pro | Gemini 3 Pro Image | gemini-3-pro-image | Complex design, world knowledge, branding, and precise control |
| Nano Banana | Gemini 2.5 Flash Image | gemini-2.5-flash-image | The first-generation old version; officials recommend migrating to it. |
Google has released Gemini-3-Pro-Image as a stable version; the old \"Preview\" designation should no longer be used for new projects. Developers should also pay attention to the release notes to avoid using consumer-facing nicknames as fixed API identifiers.
Main functions
- Generate high-fidelity images based on complex text requirements.
- Edit and upload the image, keeping the specified people, products, or layout.
- Generate clearer, more readable multilingual text in images.
- Create posters, logo drafts, packaging, and product prototypes.
- Convert data, explanations, and knowledge into infographics and diagrams.
- Search on Google to obtain information about the real world.
- Combine multiple reference images while maintaining consistency in the people and objects.
- Locally modify objects, camera angles, depth of field, lighting, and colors.
- Translate or localize existing visual content into other languages.
- Choose from various aspect ratios as well as 1K, 2K, and 4K outputs.
- Interleave text and images in a single response.
- It can be used through the Gemini app, AI Studio, APIs, and Vertex AI.
Text to image
Nano Banana Pro is suitable for complex prompts that require the model to first understand the task before planning the visuals. Compared to models that focus solely on generating results quickly, it places more emphasis on compositional logic, world knowledge, and precise control.
Gemini app creation tutorial
- Open the Gemini app and select the option to create images.
- In the model options, select Thinking or the appropriate professional image mode.
- Describe the subject, purpose, audience, composition, text, and output ratio.
- To keep a character or product, upload legitimate reference images.
- After creation, check the text, facts, identity characteristics, and brand elements.
- Continue the dialogue by modifying only one area or one design variable.
- Download the results and carry out the final proofreading in professional software.
Professional prompt structure
- Tasks: posters, infographics, product prototypes, storyboards, or advertising images.
- Audience: age, region, language, and usage environment.
- Content: Objects, data, titles, and descriptions that must be present.
- Composition: image proportions, subject position, negative space, and visual hierarchy.
- Lenses: focal length, angle, shot size, depth of field, and camera movement.
- Light: time, direction, intensity, color temperature, and shadows.
- Brand: Logo location, color codes, font preferences, and elements to be excluded.
- Output: resolution, language, transparency requirements, and delivery channels.
Text and multilingual localization in images
Nano Banana Pro improves the readability of text within images, making it suitable for posters, menus, packaging, charts, and interface prototypes. The developers still recommend determining the exact text first before asking the model to insert it into the image.
- First, generate and review the title, text, and numbers separately.
- Control the amount of text to avoid a page filled with small characters.
- Specify the language, case, punctuation, and line breaks.
- When translating images, it is required that other elements remain unchanged.
- The use of brand-specific fonts requires proper authorization; AI-generated fonts that resemble them are not the authentic versions.
- The final draft must be checked word by word; it is not sufficient to rely only on thumbnails.
Infographics and real-world knowledge
The model can convert recipes, teaching materials, data, and instructions into infographics; it can also retrieve information from Google searches to enrich the content with real-world data. Search-based information enhances timeliness, but it does not guarantee that all facts and visual representations are entirely accurate.
- First, provide the verified core data and definitions.
- It is required to indicate the unit, date, region, and statistical methodology.
- Have the model first produce a text draft, and then generate an infographic.
- Time of generation for weather, sports, and market data records.
- Medical, legal, financial, and research-related information must be reviewed manually.
- It is difficult to include the full source information in the final image; a separate list of references should be attached.
Multiple reference figures and consistency
According to the official documentation, the Nano Banana Pro can process up to 14 input images, of which 5 can be retained with high fidelity. The initial product description also states that it is possible to maintain the similarity of up to 5 individuals.
| Input scale | Official capability description | Usage suggestions |
|---|---|---|
| 1 to 5 pieces | Supports high-fidelity references | Suitable for materials related to people, products, and brands |
| Up to 14 pieces | More visual elements can be combined. | Consistency may decline in complex combinations. |
| Up to 5 characters | Emphasize the preservation of character similarities. | It is still necessary to check the face shape, age, clothing, and hands. |
| Various objects | It can be used for building product portfolios and scenarios. | Specify the attributes that should be retained for each reference image. |
Refer to the diagrammatic operation steps for more details.
- Only reference images that have the rights to be uploaded and generated should be selected.
- Number each image and explain its purpose.
- Distinguish between content that must be retained and content that can vary freely.
- First, use a small number of reference images to test the consistency of the characters and products.
- Gradually add elements such as backgrounds, clothing, accessories, and brand elements.
- In each round, only one type of variable is adjusted, and a stable version is saved.
- Zoom in to examine the face, hands, text, logos, and product structure.
Local editing and creative control
Nano Banana Pro allows for the modification of specific elements through dialogue, and it enables control over the camera angle, focus, depth of field, lighting, and color grading. Generative editing involves re-estimating pixels, so it cannot guarantee that areas that are not specified will remain unchanged.
- Change the daytime scene to night while keeping the subject intact.
- Change the direction of the light source, the contrast between light and dark, and the film’s color tone.
- Adjust the focus and depth of field to highlight a specific object.
- Expand or shrink the background to fit the new screen ratio.
- Replace a single element in the costumes, props, and scenery.
- Convert sketches, blueprints, or conceptual plans into realistic representations.
- Specify clear locking requirements for areas that do not require any changes.
Output resolution and scale
The Nano Banana Pro supports 1K, 2K, and up to 4K images, offering a variety of common aspect ratios. 4K output is more costly, and it does not automatically correct all issues related to text, textures, and structure.
| Output | Suitable uses | API standard output price |
|---|---|---|
| 1K | Web drafts, social media images, and quick review | Approximately 0.134 dollars per sheet |
| 2K | Higher-definition design, presentations, and medium-sized outputs | Approximately 0.134 dollars per sheet |
| 4K | High-resolution final outputs, pre-printing materials, and detailed presentations | Approximately 0.24 dollars per sheet |
The prices mentioned above cover only the costs associated with the standard API for image generation; they may also include expenses related to text input, image input, processing time, and search functions. The Gemini plans for end-users are not billed based on this API pricing table.
Interleaved text and image output
Gemini 3 Pro Image can generate text paragraphs and illustrations interleaved within the same response, making it suitable for stories, instructional guides, and step-by-step instructions. Developers need to iterate through the content blocks in the response, rather than only reading a single image field.
- Illustrations corresponding to the story content can be inserted between chapters.
- The teaching steps can provide both textual explanations and schematic diagrams at the same time.
- Complex responses need to be saved and rendered in the order of their content blocks.
- The frontend should handle exceptions such as missing images, duplicate images, and empty text.
- The cost of generating each output image is taken into account.
Entry point for use in the Gemini app
Nano Banana Pro can be used on the Gemini web interface as well as in its mobile app; access is usually obtained by creating an image and selecting a Thinking-class model. Free users have limited quotas, while Google AI Plus, Pro, and Ultra users enjoy higher quotas.
| Entrance | Available methods | Explanation |
|---|---|---|
| Gemini app | Create an image and select the Thinking model. | The free quota is limited, while paid subscriptions offer a higher quota. |
| Google AI Mode | Professional image mode is used in some countries and languages. | There are many restrictions regarding region, account, and plan options. |
| NotebookLM | Create infographics and slide decks | Visualization can be based on user profiles and research results. |
| Google Slides | Help me visualize and beautify slides. | For eligible Workspace customers |
| Google Vids | Generate image materials for video projects | For eligible Workspace customers |
| Flow | Generating and controlling video footage | Nano Banana Pro is aimed at paid plans. |
| Mixboard | Canvas and presentation capabilities | They are experimental products from Google Labs. |
Watermarks and AI content identification
All outputs from Google’s image generation models contain an invisible SynthID digital watermark. Images generated for free within the Gemini app, as well as those created using Google AI Pro, may also have visible Gemini markers; meanwhile, the visible watermarks can be removed from outputs generated by Ultra and AI Studio.
- The invisible SynthID is used to identify content generated by Google AI.
- Removing visible markers does not mean that the SynthID has been removed.
- Platforms and regions may require additional disclosure regarding AI-generated content.
- Ads, news, and public communications should clearly indicate synthetic content.
- Do not use AI-generated images to pretend to be real people, events, or evidence.
Google AI Personal Plan
Google One currently offers three personal AI plans: Google AI Plus, Pro, and Ultra. Prices, taxes, additional storage space, and product quotas may vary depending on the region. The main text of the catalog should not state a fixed price in RMB that applies everywhere in the world.
| Package | Benefits related to Nano Banana Pro | Suitable for users |
|---|---|---|
| Free tier | The Gemini application has a limited quota; once this limit is reached, another image model must be used. | Occasional generation related to functional experience |
| Google AI Plus | Image and AI quotas at a level higher than the free tier | Light creation and everyday use |
| Google AI Pro | Higher usage limits for Nano Banana Pro and AI Studio | Students, creators, and professionals |
| Google AI Ultra | Maximum quota, more Flow capacity, and professional watermark features | High-frequency creation and advanced experimental features |
The actual price of the plans and the availability of their features are determined by the purchase page specific to the country or region where the Google One account is located. Workspace, school, and enterprise accounts are also subject to the permissions set by administrators as well as the constraints imposed by the organizational plans.
Gemini API technical specifications
| Project | Gemini 3 Pro Image |
|---|---|
| Model code | gemini-3-pro-image |
| Input type | Text and images |
| Output type | Text and images |
| Maximum input tokens | 65536 |
| Maximum output tokens | 32768 |
| Image generation | Support |
| Reasoning/Thinking | Support |
| Google search grounding | Support |
| Batch API | Support |
| Flex reasoning | Support |
| Priority reasoning | Support |
| Function Calling | Not supported |
| Structured Output | Not supported |
| URL Context | Not supported |
| File search and caching | Not supported |
| Audio and Live API | Not supported |
Standard price for Gemini API
The following are the prices in US dollars as of August 2026, based on the pricing page for the Google Gemini Developer API. The Nano Banana Pro does not offer a free tier for this API; paid accounts are required for any usage.
| Billing items | Standard paid price | Conversion notes |
|---|---|---|
| Text or image input | 2 dollars per million tokens | Each individual image input is counted as 560 tokens, which is approximately 0.0011 dollars. |
| Text and thought output | 12 dollars per million tokens | Charged based on the actual number of tokens generated. |
| Image output | 120 dollars per million tokens | 1K and 2K cost around $0.134, while 4K costs around $0.24. |
| Google search grounding | 5,000 free search requests per month, after that 14 dollars per 1,000 requests | The Gemini 3.x models share a free quota, which is calculated based on the actual number of search queries. |
Prices for Batch, Flex, and Priority
| Pattern | Enter price | Text and thought output | Image output | Suitable for tasks |
|---|---|---|---|---|
| Batch | Text: 1 dollar per million tokens; approximately 0.0006 dollars per image. | 6 dollars per million tokens | 1K or 2K is about 0.067 dollars; 4K is about 0.12 dollars. | Batch generation without real-time processing |
| Flex | Text: 1 dollar per million tokens; approximately 0.0006 dollars per image. | 6 dollars per million tokens | 1K or 2K is about 0.067 dollars; 4K is about 0.12 dollars. | Tasks with tolerable elastic capacity |
| Priority | Text or images: 3.60 dollars per million tokens | 21.60 dollars per million tokens | 216 dollars per million image tokens | Production operations that require priority in terms of capacity and low latency |
The discounts for Batch and Flex are significant, but their response times and capacity guarantees differ. Companies can also make purchases through Vertex AI; costs related to location, committed usage levels, networking, and other cloud services need to be calculated separately.
API Integration Tutorial
- Create a project in Google AI Studio and enable the paid Gemini API.
- Create an API Key and store it only in the backend key management environment.
- Install the official Google Gen AI SDK and specify a compatible version.
- Send the minimum request using the stable model code gemini-3-pro-image.
- Configure the output image ratio, size, and response mode.
- Save the text and image content blocks in the response.
- Record the actual costs of input, output, grounding, and retry attempts in case of failures.
- Add rate limiting, timeouts, exponential backoff, and content filtering.
- Pre-production testing areas, quotas, delays, and data governance.
API cost control
- For creative exploration, start by using Nano Banana 2 Lite or 2.
- Only complex, high-value pages should be handed over to Pro for generation.
- 1K is used during the preview phase, while 4K is generated after finalization.
- Reduce duplicate reference figures and unnecessary long prompt texts.
- The search for a ground connection is enabled only when real-time information is truly necessary.
- Offline tasks use Batch or Flex to reduce the cost of image output.
- Set daily budgets, quotas, and anomaly alerts for users and projects.
How to choose between Nano Banana Pro and Series 2
| Demand | Recommended models | Reason |
|---|---|---|
| High-throughput rapid sketching | Nano Banana 2 Lite | Speed and cost are prioritized. |
| Universal generation and rapid multi-round editing | Nano Banana 2 | Balance between quality, 4K, text, and speed |
| Complex information diagrams and world knowledge | Nano Banana Pro | Reasoning and search have a stronger basis in reality. |
| High fidelity for brands and multiple reference images | Nano Banana Pro | Precise control and consistency take precedence. |
| Mass final output | First, conduct a screening using Series 2, and then finalize it with Pro. | Concentrate the high-cost models on the key finished products. |
| Existing projects from the old version | Migrate to 2 Lite or 2 | The authorities have classified the first generation as an older version. |
What use cases are suitable?
- Designers create high-fidelity posters, packaging, and product prototypes.
- The brand team combines products, characters, and visual guidelines.
- Educators convert knowledge and materials into infographics.
- The marketing team creates multilingual ads and localized content.
- The film and television team creates storyboards, scene designs, and lighting concept art.
- The e-commerce team creates product scenarios and different market versions.
- Developers create applications for image generation, editing, and design.
- Workspace users create visual content in Slides and Vids.
- NotebookLM users convert source materials into presentations and illustrations.
Tasks that are not suitable for completion directly
- It is required that all facts and data in the image be absolutely accurate.
- Precise font files, vector layers, and print color management are required.
- Generating celebrities, customers, or protected characters without authorization.
- Use AI-generated images as evidence in news, medical, or legal matters.
- Projects that must be completely offline and whose assets cannot be sent to the cloud.
- A production pipeline that relies on fixed random outcomes and pixel-level reproducibility.
- Use images from the free Gemini app for direct commercial use.
- It requires the model to perform function calls or output strictly structured data.
Product advantages
- Gemini’s reasoning capabilities and its knowledge of the world are suitable for complex visual tasks.
- The text within images, along with the ability for multilingual localization, are quite outstanding.
- It supports generating visual content related to real-world information through Google searches.
- Up to 14 reference images can be processed.
- Provides greater control over the consistency of characters, products, and brands.
- It supports local editing, as well as adjustments to the lens, depth of field, lighting, and color.
- It offers 1K, 2K, and up to 4K output.
- Text and images can be returned alternately.
- It covers Gemini, NotebookLM, Workspace, AI Studio, and Vertex AI.
- The stable API model supports Standard, Batch, Flex, and Priority modes.
Usage restrictions and precautions
- The Nano Banana Pro is not the fastest or cheapest model available at the moment.
- There is no free tier for the API; costs are incurred for input, output, and grounding.
- The price for 4K single-image output is higher than that for 1K and 2K.
- The model does not necessarily return the specified number of images as required.
- Up to 14 inputs do not mean that each one can maintain the same level of fidelity.
- The text in the images may still contain spelling errors or an improper layout.
- Searching for a ground point may result in multiple queries, with each one being charged separately.
- The model does not support audio input, function calls, or structured output.
- Consumption quotas, entry points, and watermarks vary depending on the package and region.
- The generated results may contain factual errors, biases, or inappropriate elements.
- The user is responsible for the rights to third-party trademarks, characters, figures, and reference images.
Privacy and data security
The Gemini API pricing page states that the data from the Nano Banana Pro subscription tier is not used to improve Google’s products. Gemini for consumers, Workspace, and Vertex AI are subject to different data policies, and the API subscription protection cannot be applied to all of these platforms.
- Sensitive operations should prioritize the use of enterprise contracts and controlled Vertex AI environments.
- Do not put the API Key in browsers, mobile applications, or public code.
- Obtain explicit authorization before uploading images of people, products, and customers.
- Remove location, device, and personal identification information from the images.
- Record the reference image, prompt, model version, and generation time.
- Set minimum permissions, budgets, and log access controls for the team.
- Verify the rules for data residency, retention, and deletion based on the business region.
GitHub and the open-source status
The weights, training data, and Google-hosted services related to the Nano Banana Pro model are not made available under an open-source license. Google has published its official Google Gen AI Python SDK and Gemini Cookbook on GitHub, both of which are licensed under the Apache 2.0 license.
The fact that the SDKs and examples are available under an open-source license does not mean that the models are also open source; moreover, it is not allowed to run Gemini 3 Pro Image locally, outside of Google’s services. Developers still need an API key, a paid account, and to make use of cloud-based services in compliance with Google’s terms.
Basic information
| field | Content |
|---|---|
| Product name | Nano Banana Pro |
| Official model name | Gemini 3 Pro Image |
| Developer | Google DeepMind |
| Stable API code | gemini-3-pro-image |
| Tool type | AI image generation, editing, and multimodal visual models |
| Enter | Text and images |
| Output | Text and images |
| Highest resolution | 4K |
| Refer to the diagram | Up to 14 cards, of which 5 are in high-fidelity quality. |
| Search for grounding | Supports Google search |
| Main entrance | Gemini, AI Studio, Gemini API, Vertex AI, and Google products |
| Whether API is provided | Yes, the payment tiers. |
| Is it open source? | The model is not open-source; the official SDK and Cookbook use Apache 2.0 |
| AI content identification | All generated images contain SynthID. |
Recommendation score
4.8 / 5. The Nano Banana Pro boasts comprehensive capabilities in complex visual reasoning, multilingual text handling, information graphics, use of multiple reference images, and brand control, and it offers a stable API; however, cost, fact verification, permissions, as well as watermarking and copyright management remain essential aspects for professional use.
Frequently Asked Questions
What is the official name of Nano Banana Pro?
Its official name is Gemini 3 Pro Image, while the code for its stable API model is gemini-3-pro-image. Nano Banana Pro is the nickname used by Google for this product.
Is the Nano Banana Pro still the latest model?
It remains the high-end model in the family designed for the most complex professional tasks, but Google has introduced the faster Nano Banana 2 and 2 Lite. Different models offer different trade-offs in terms of quality, speed, and cost.
Can it be used for free?
The Gemini app offers free users a limited amount of Nano Banana Pro credits; once that limit is reached, other image models are used. The Gemini API does not provide a free version of this model.
How much does it cost to generate an image with API?
In the standard API, the cost for outputting images in 1K or 2K resolution is around $0.134, while the cost for 4K resolution is around $0.24; in addition, costs related to input data, text and thought outputs, as well as possible search-related fees, must also be taken into account.
How many reference images can be uploaded at most?
The official documentation specifies a maximum of 14 input images, 5 of which support higher fidelity. For complex combinations, manual inspection is still required to ensure the consistency of characters and objects.
Is it possible to generate text in Chinese?
It supports multiple languages including Chinese, and the functionality for handling text within images as well as for localization has been improved. The final version still needs to be proofread word by word.
Are there watermarks on the generated images?
All outputs contain an invisible SynthID. Some packages for the Gemini consumer version include visible markers, while options such as Ultra and AI Studio allow for output without any visible watermarks.
Is Nano Banana Pro open source?
The model is not open source. Google’s official SDK and Cookbook are available under an open source license based on Apache 2.0, but it is still necessary to use the models hosted by Google.
Guigong Network Security Registration No. 45132202000164