Cloudglue
Free value-added services
Comprehensive List of AI Tools AI programming tools

Cloudglue

Cloudglue makes AI programming more efficient and simpler.

Tags:

What is Cloudglue?

Cloudglue is a video and audio context platform developed by Aviary Inc.; its core components are a web-based console, development APIs, and software development kits. It organizes voice, speakers, visuals, on-screen text, audio, and metadata into a context that can be searched, extracted, and used for inference by large models.

This product is designed to serve developers and teams that need to process large amounts of media content within their applications, rather than being a traditional video editor that offers functions such as timeline editing, filters, or video generation.

Main functions

Multimodal video description

  • Describe can generate video descriptions in chronological order; the results may include summaries, transcripts, identification of speakers, visual scenes, on-screen text, and audio events.
  • Users can select the desired mode to reduce unnecessary processing; the resulting output is suitable for subtitle preparation, content review, accessible description, and search indexing.
  • The description of the results can be read in segments, and thumbnails can be included to facilitate navigation to the corresponding moments in the video.

Structured data extraction

  • Extract accepts natural language prompts or custom structures, and outputs entities and fields that can be further processed by programs.
  • The hint mode is suitable for exploring what information is available in a video, while the structure mode is appropriate for repeatedly extracting the same set of fields from interviews, product videos, sales calls, or research materials.
  • Entity sets allow unified prompts and structures to be applied to a set of videos, helping teams create comparable datasets.

Segments, shots, and chapters

  • Segment allows media to be divided into shots or narrative chapters, returning the start and end times as well as a description for each segment.
  • Developers can use the segmentation results to generate chapter navigation, candidate clips, keyframe indices, or fine-grained search units.
  • Narrative segments designed for transcription are suitable for meetings and podcasts that rely primarily on spoken content; video segments are more appropriate for videos with significant visual changes.

Cross-video search and Q&A

  • Search supports semantic retrieval using natural language, enabling the identification of relevant videos, clips, and specific moments within files or collections.
  • Mixed retrieval allows for the combination of various modalities such as regular text, voice keywords, on-screen text, and tags, which helps to avoid missing visual information when only transcriptions are searched.
  • The Chat Completions and Responses API allows for sequential queries based on one or more collections, returning media references accompanied by timestamps.
  • Deep Search carries out multiple steps of retrieval and reasoning, making it suitable for identifying research answers by examining a large number of videos; preview models should not be regarded as reliable guarantees.

Collection and Media Management

  • Collections are used to organize media related to the same project, and indexes for various purposes such as media descriptions and entities can be created.
  • The file interface allows one to view the processing status, duration, resolution, encoding method, and thumbnails; it also supports functions such as tagging, sharing assets, frame extraction, as well as face detection and matching.
  • Webhooks can notify the business system when asynchronous processing is completed or when there is a change in status, thereby avoiding continuous polling by the client.

Differences in input methods and capabilities

Input methodSupported contentSuitable usesImportant restrictions
Local uploadMP4, MOV, WEBM; common audio files are also supported.Complete multimodal description, extraction, search, and question answeringLarge files and batch tasks are better suited for connectors or interfaces.
Public linkDirectly accessible media addresses, TikTok, LoomQuick import of public assetsThe link must be accessible to the service, and the user remains responsible for the platform permissions.
YouTubePublic video linkVoice, transcription, and metadata tasksCurrently, only audio and metadata are processed; for full visual understanding, the original files must be uploaded.
Cloud connectorGoogle Drive, Dropbox, Zoom, Gong, Grain, S3, GCS, Recall.ai, iconikBatch import and production pipelineSeparate authorizations are required; terminating the connection does not automatically delete the imported files.

Typical usage process

  1. Register an account and create an API key in the console; the key should be stored as a server-side environment variable, and not included in the frontend code or in public repositories.
  2. Upload local media, paste a public link, or authorize a cloud connector, then check the status of file processing.
  3. For single-file tasks, it is possible to directly use the Describe, Extract, or Segment functions; for batch retrieval and question-answering, it is first necessary to create a collection and add the files to it.
  4. Select a summary, audio, video, on-screen text, structural fields, or segmentation strategy based on the task, and configure status queries or Webhooks for asynchronous tasks.
  5. Use Search, Chat, or Responses to query the collection; the timestamp and media references returned are then mapped to the player or the business interface.
  6. Before going live, record the score consumption, response status, and retry attempts for each endpoint, and set the concurrency level and budget based on the actual limits of the account.

Suitable for users and scenarios

  • Development team: Add video search and Q&A capabilities to knowledge bases, customer service systems, or chatbots, without the need to manage transcription, visual modeling, and vector search pipelines themselves.
  • Sales and Customer Success Teams: Extract objections, requirements, product feedback, and action items from calls and demonstrations, while retaining identifiable excerpts.
  • Research and Media Team: Locate instances of individuals, topics, visuals, and original quotes in interviews, courses, podcasts, or resource libraries.
  • Platform-based products: they provide structured media data for video management, content moderation, meeting analytics, or educational applications.
  • Automation team: Utilizes connectors, interfaces, Webhooks, and MCP to integrate media analysis into existing workflows or AI assistants.

Price packages

The public prices are displayed in dollars as package of points; for subscription-based needs and solutions for large teams, it is necessary to contact sales. The page does not specify the validity period of the point packages, taxes, rules regarding refunds, or what happens to unused points – these details should be checked on the payment page before making a purchase.

Package or versionPriceBilling cycleCore benefits or quotaSuitable for users
Free0 dollarsFree quota200 points, 200 chat requests, 25 minutes of indexing timeExperience and small prototypes
Mini15 dollarsPoints package; the period is not specified.1000 points, 1000 chat requests, 2 hours of indexing timePersonal testing and lightweight projects
Starter45 dollarsPoints package; the period is not specified.3000 points, 3000 chat requests, 6 hours of indexing timeContinuously developed small applications
Builder350 dollarsPoints package; the period is not specified.30,000 points, 30,000 chat requests, 63 hours of indexing timeHigher volumes for production and development
Scale600 dollarsPoints package; the period is not specified.60,000 points, 60,000 chat requests, 125 hours of indexingMass media processing
EnterpriseContact salesCustomizationCustom solutions and discounts for large teamsCorporate procurement and special capacity requirements

The terms of service specify that the subscription is on a monthly pay-as-you-go basis, with invoices issued on a monthly basis and payment required within 5 days after receipt of the invoice. This description differs slightly from what is stated on the public page under “Select a points package”; the actual method of purchase, as well as the conditions for renewal and payment, should be based on the account settlement page or the corporate order details.

Main points consumption

OperationPublic scoring rulesMeasurement method
Extract4 pointsMaterials per minute
Transcribe or Describe4 pointsMaterials per minute
Upload videoIncludesFile operations
Entity set index4 pointsMaterials per minute
Media collection index4 pointsMaterials per minute
Chat Completion1 pointEach request
Search1 pointEach search

Points are charged only when a request is successful, based on the usage per endpoint and function. For certain items such as scenario segmentation, the values listed in the public tables are not clear enough to conclude that they are free of charge.

Limits and usage restrictions

  • The interface enforces rate and usage limits, and errors are returned if these limits are exceeded; for higher limits, it is necessary to contact the team.
  • The quota documentation still uses the Free Tier and Basic Tier, while the current pricing page uses names such as Mini, Starter, Builder, Scale, etc.; there is no automatic correspondence between the two.
  • On the quota page, both the Free and Basic plans allow for a maximum of 100 files per month, a total storage capacity of 10GB, and a total video duration of 1000 hours. The number of collections allowed is 50 for the Free plan and 100 for the Basic plan; the number of files in each collection is 50 for the Free plan and 100 for the Basic plan.
  • The same page also indicates that the number of withdrawals and transcriptions per month is 200 and 400 respectively; the chat quota is 500 per month and 50 per day for the Free plan, and 30,000 per month and 1,000 per day for the Basic plan.
  • These technical quotas, along with the points on the price page and the number of chat requests, are all relevant factors; it is necessary to check the actual limits of the current package in the account dashboard before proceeding with production deployment.

API, SDK, and MCP

ComponentsAccess methodLicenseExplanation
Cloudglue APIREST and Bearer API keysProprietary servicesCovering files, collections, descriptions, extraction, searching, chatting, Responses, and management endpoints
JavaScript SDKnpm packagesApache-2.0Suitable for Node.js and TypeScript applications
Python SDKPython packagesApache-2.0Suitable for data processing, backend tasks, and automation scripts
MCP Servernpm commands or desktop extensionsMITThe video collection tool can be integrated into AI clients that are compatible with MCP.

What is open source are the SDK and MCP Server client code; this does not mean that Cloudglue’s hosting platform, model orchestration capabilities, or infrastructure are open source. To use these components, one still needs a Cloudglue account, keys, and credits, and they are subject to the platform’s terms of service.

Capacity boundaries

  • Automatic description, transcription, entity extraction, and reasoning may miss details or produce errors; in legal, medical, security, and compliance contexts, manual review of the original media is necessary.
  • YouTube links currently lack full visual understanding; if the task relies on actions, objects, text on the screen, or changes in camera angle, the original file that can be processed should be uploaded.
  • Video processing is typically an asynchronous task, and the latency is influenced by duration, resolution, the selected mode, as well as the size of the queue and collections.
  • Deep Search and the preview models are in a phase of rapid development; their interface parameters, quality, and costs may change. In a production environment, it is necessary to stick to a specific version and review the update logs.
  • Playing the referenced material can help in locating evidence, but it does not guarantee that the information is accurate; the final judgment should still be based on reviewing the relevant footage.

Privacy, security, and copyright

  • The service handles registered email addresses, third-party login credentials, uploaded videos or data, as well as analysis of website and product usage.
  • The privacy policy states that PostHog is used for behavior analysis and Stripe is used for payments, and it says that personal information will not be sold, rented out, or traded.
  • Personal information and uploaded content are retained for as long as is necessary to provide services, fulfill legal obligations, and resolve disputes; however, no specific period for retention or automatic deletion has been specified.
  • Users in the European Economic Area can request access to, correction of, deletion of, porting of, restriction on, or objection to processing; deleting a connector does not automatically delete the media that has been imported previously.
  • Users retain the right to the videos and data they upload, and must confirm that they have the authorization to process, transcribe, analyze, and share such materials.
  • The terms prohibit the handling of illegal, harmful, or pornographic content, as well as bypassing usage limits, conducting reverse engineering, reselling the interfaces without written permission, and using the platform to create or operate competitive products.
  • Copying, modifying, and redistributing content generated by the platform is permitted only when explicitly allowed by the terms or the relevant SDK license; before deploying such content in a commercial setting, it is necessary to check the copyright status of the material, privacy permissions, and any applicable customer contracts.

Refunds and purchasing considerations

The public terms do not specify in detail the rules regarding refunds for self-service point packages, cancellation of purchases, or the return of remaining points. Corporate orders, monthly invoicing, and self-service point packages may be subject to different conditions; it is necessary to save the settlement page before making a payment and to check the details related to refunds, expiration dates, renewals, and taxes.

Summary

The value of Cloudglue lies in its ability to combine multi-modal video analysis, collection indexing, retrieval, question answering, and structured data extraction into a single development interface, while also providing SDKs, MCPs, and various media connectors. It is suitable for product teams that need to transform videos into queryable business data; however, before going live it is essential to verify the limitations related to YouTube content, the current account limits, the costs associated with processing, and compliance requirements regarding content handling.

©️Copyright notice: Unless otherwise specified, all articles on this site are copyrighted bySharing of AI toolsAll content on this site is original; without permission, no individual, media outlet, website, or organization may reproduce, copy, or otherwise distribute it, nor may they create mirrors of it on servers that are not owned by this site. Otherwise, we reserve the right to take legal action against such parties in accordance with the law.

Tools similar to Cloudglue