Batch Processing
capability · v1.0.0 Supports batched inference requests processed asynchronously at reduced cost, with results retrieved after processing completes.
84 atoms across 2 types.
capability · v1.0.0 Supports batched inference requests processed asynchronously at reduced cost, with results retrieved after processing completes.
capability · v1.0.0 Ability to invoke structured tool definitions from model output. Enables agentic workflows where the model selects and calls tools based on context.
capability · v1.0.0 Supports inference with Low-Rank Adaptation (LoRA) fine-tuning adapters overlaid on base model weights at inference time.
capability · v1.0.0 Supports streaming output or real-time bidirectional interaction with low latency, suitable for voice agents and interactive UIs.
capability · v1.0.0 Enhanced chain-of-thought reasoning that produces step-by-step thinking before a final answer, improving accuracy on complex tasks.
capability · v1.0.0 Guarantees model output conforms to a caller-supplied JSON Schema or equivalent structure, stronger than JSON mode alone.
capability · v1.0.0 Ability to process image inputs alongside text for visual recognition, image reasoning, and captioning.
model-card · v1.0.0 First open-source transformer-based multilingual NMT model supporting high-quality translation across all 22 scheduled Indic languages.
model-card · v1.0.0 Southeast Asian Languages In One Network — LLMs pretrained and instruct-tuned for the Southeast Asia region.
model-card · v1.0.0 General embedding base model transforming text into 768-dimensional vectors.
model-card · v1.0.0 General embedding large model transforming text into 1024-dimensional vectors.
model-card · v1.0.0 Multi-functionality, multi-linguality, and multi-granularity embedding model.
model-card · v1.0.0 Reranker model that takes question and document as input and outputs a relevance similarity score.
model-card · v1.0.0 General embedding small model transforming text into 384-dimensional vectors.
model-card · v1.0.0 12 billion parameter rectified flow transformer capable of generating images from text descriptions.
model-card · v1.0.0 Image model capable of generating highly realistic and detailed images with multi-reference support.
model-card · v1.0.0 Ultra-fast distilled image model delivering state-of-the-art quality for interactive workflows and real-time previews.
model-card · v1.0.0 Ultra-fast distilled image model with enhanced quality; unifies image generation and editing in a single model.
model-card · v1.0.0 Lightning-fast text-to-image generation model capable of producing high-quality 1024px images in a few steps.
model-card · v1.0.0 Context-aware text-to-speech model applying natural pacing and expressiveness.
model-card · v1.0.0 Context-aware English text-to-speech model that applies natural pacing, expressiveness, and fillers based on context.
model-card · v1.0.0 Context-aware Spanish text-to-speech model that applies natural pacing, expressiveness, and fillers based on context.
model-card · v1.0.0 First conversational speech recognition model built specifically for voice agents.
model-card · v1.0.0 Deepgram speech-to-text model for transcribing audio.
model-card · v1.0.0 SQL-specialized model designed to help non-technical users understand and query data in SQL databases.
model-card · v1.0.0 300M parameter state-of-the-art open embedding model producing vector representations for search and retrieval; trained with data in 100+ languages.
model-card · v1.0.0 Gemma 2B base model dedicated for inference with LoRA adapters.
model-card · v1.0.0 128K context window model with multilingual support in over 140 languages; handles text and image input.
model-card · v1.0.0 Google's most intelligent family of open models, derived from Gemini 3 research.
model-card · v1.0.0 Lightweight, state-of-the-art open model built from the same research and technology as Gemini models.
model-card · v1.0.0 Gemma 7B base model dedicated for inference with LoRA adapters.
model-card · v1.0.0 Distilled BERT model fine-tuned on SST-2 for sentiment classification.
model-card · v1.0.0 Industry-leading results in agentic tasks including instruction following and function calling; suited for RAG, multi-agent workflows, and edge deployments.
model-card · v1.0.0 Most adaptable and prompt-responsive model with strengths in sharp graphic design, full-HD renders, and accurate text rendering.
model-card · v1.0.0 Generates images with exceptional prompt adherence and coherent text rendering.
model-card · v1.0.0 Open-source multimodal chatbot trained by fine-tuning LLaMA/Vicuna on visual instruction data; supports image captioning and visual question answering.
model-card · v1.0.0 Stable Diffusion model fine-tuned to be better at photorealism.
model-card · v1.0.0 Distilled from DeepSeek-R1 based on Qwen2.5; outperforms OpenAI o1-mini on several benchmarks.
model-card · v1.0.0 DETection TRansformer model trained end-to-end on COCO 2017 object detection dataset (118k annotated images).
model-card · v1.0.0 Full precision (fp16) generative text model with 7 billion parameters.
model-card · v1.0.0 Llama 2 base model dedicated for inference with LoRA adapters.
model-card · v1.0.0 Quantized (int8) generative text model with 7 billion parameters.
model-card · v1.0.0 State-of-the-art performance on industry benchmarks with improved reasoning capabilities.
model-card · v1.0.0 Quantized (int4) generative text model with 8 billion parameters.
model-card · v1.0.0 Large multilingual model optimized for multilingual dialogue use cases.
model-card · v1.0.0 Multilingual large language model optimized for multilingual dialogue use cases.
model-card · v1.0.0 Quantized (int4) generative text model with 8 billion parameters.
model-card · v1.0.0 Fast variant of Llama 3.1 8B optimized for multilingual dialogue use cases.
model-card · v1.0.0 Llama 3.1 8B quantized to FP8 precision.
model-card · v1.0.0 Multimodal model optimized for visual recognition, image reasoning, and captioning.
model-card · v1.0.0 Lightweight model optimized for multilingual dialogue use cases.
model-card · v1.0.0 Lightweight model optimized for multilingual dialogue including agentic retrieval and summarization.
model-card · v1.0.0 Llama 3.3 70B quantized to fp8 precision and optimized for faster inference.
model-card · v1.0.0 17 billion parameter natively multimodal model with 16 experts (mixture-of-experts architecture).
model-card · v1.0.0 Llama 3.1 8B fine-tuned for content safety classification in both LLM inputs and responses.
model-card · v1.0.0 Multilingual encoder-decoder model trained for many-to-many multilingual translation.
model-card · v1.0.0 State-of-the-art performance on a wide range of industry benchmarks.
model-card · v1.0.0 Transformer-based model with next-word prediction objective, trained on 1.4T tokens from web and synthetic sources.
model-card · v1.0.0 50 layers deep image classification CNN trained on more than 1 million images from the ImageNet dataset.
model-card · v1.0.0 Instruct fine-tuned version of the Mistral-7B model.
model-card · v1.0.0 32k context window instruct model with updated rope-theta and no sliding-window attention.
model-card · v1.0.0 Mistral 7B instruct v0.2 dedicated for inference with LoRA adapters.
model-card · v1.0.0 State-of-the-art vision understanding with long context capabilities up to 128k tokens.
model-card · v1.0.0 Frontier-scale open-source model with a 256k context window.
model-card · v1.0.0 Frontier-scale open-source 1T parameter model with a 262.1k context window, designed for agentic workloads.
model-card · v1.0.0 High-quality multi-lingual text-to-speech library.
model-card · v1.0.0 Upgraded, retrained version of Nous Hermes 2 with function calling and JSON mode capabilities.
model-card · v1.0.0 Hybrid MoE model with leading accuracy for multi-agent applications.
model-card · v1.0.0 Open-weight model designed for production, general-purpose, high-reasoning use cases.
model-card · v1.0.0 Open-weight model optimized for lower latency and local or specialized use cases.
model-card · v1.0.0 General-purpose speech recognition model supporting multilingual recognition, speech translation, and language identification.
model-card · v1.0.0 Pre-trained model for automatic speech recognition (ASR) and speech translation.
model-card · v1.0.0 English-only version of the Whisper Tiny model trained on speech recognition.
model-card · v1.0.0 Japanese text embedding model that converts Japanese text input into numerical vectors for information retrieval, text classification, and clustering.
model-card · v1.0.0 Open source community-driven native audio turn detection model in its second version.
model-card · v1.0.0 Code-specific large language model supporting six mainstream model sizes from 0.5B to 32B parameters.
model-card · v1.0.0 Latest generation large language model with groundbreaking advancements in reasoning, instruction-following, and agent capabilities.
model-card · v1.0.0 Latest Qwen family model specifically designed for text embedding and ranking.
model-card · v1.0.0 Medium-sized reasoning model capable of thinking deeply for hard problems; competitive with DeepSeek-R1 and o1-mini.
model-card · v1.0.0 Generate a new image from an input image using Stable Diffusion.
model-card · v1.0.0 Stable Diffusion model with inpainting capability using a mask to selectively edit image regions.
model-card · v1.0.0 Diffusion-based text-to-image model that generates and modifies images based on text prompts.
model-card · v1.0.0 Small generative vision-language model primarily designed for image captioning and visual question answering.
model-card · v1.0.0 Fast and efficient multilingual text generation model with a 131,072 token context window supporting 100+ languages.