Any model. Any modality. Zero lock-in.
Draft & Goal is model-agnostic by design. Choose the best model for each step — frontier or open-weight, hosted or quantized on your own GPUs — and work across text, vision, image, audio, music, video, and reasoning in a single workflow. Swap providers without re-platforming.
Maximum flexibility, by design.
Model-agnostic
Every major provider behind one interface. Pick a model per node and swap anytime — no rewrites.
Open + closed
Frontier closed APIs or open-weight models. Choose the cost, latency, and control trade-off you want.
Truly multimodal
Text, vision, image, audio, music, and video — composed in a single governed workflow.
Reasoning on demand
Switch on extended thinking for hard steps; route fast, cheap models for everything else.
Quantized & self-hosted
Run 4-bit / 8-bit open models on your own GPUs via vLLM, Ollama, or TGI — in your VPC.
No lock-in
Bring your own keys and private endpoints. Route, fall back, and cost-optimize across vendors.
Frontier or open-weight — your call.
Connect the models your teams already trust and the new ones shipping every month (June 2026 lineup shown). Closed frontier APIs and open-weight models live side by side; pick the right one for each task.
GPT-5 · o-series · Sora · Whisper
Claude Opus 4.8 · Sonnet 4.6 · Fable 5
Gemini 3 Pro · Gemma 3 · Imagen 4 · Veo 3
Llama 4 (open weights)
Mistral Large · Magistral · Pixtral
DeepSeek-R1 · DeepSeek-V3
Qwen3 · Qwen-VL
Grok 4
Kimi K2
Not just text.
Compose across modalities in the same workflow — transcribe a call, reason over it, generate the visuals, and render a video, each on the best model for the job.
Text & language
Generation, extraction, classification, translation — GPT-5, Claude Opus 4.8, Gemini 3, Llama 4, Mistral, Qwen3.
Vision
Read images, screenshots, charts, and PDFs — GPT-5, Claude, Gemini 3, Qwen-VL, Pixtral, Llama 4.
Image generation
On-brand visuals at scale — GPT Image, Imagen 4, FLUX.1, Stable Diffusion 3.5, Midjourney v7, Ideogram.
Audio & speech
Transcription and natural voice — Whisper, OpenAI Realtime, ElevenLabs, Deepgram, Gemini audio.
Music & sound
Tracks, jingles, and sound design — Suno, Udio, Stable Audio, ElevenLabs SFX.
Video generation
Clips and product shots — Sora, Veo 3, Runway Gen-4, Kling 2, Seedance, Pika.
Reasoning
Deep multi-step thinking — o-series, DeepSeek-R1, Claude extended thinking, Gemini thinking, Magistral, Qwen3.
Run it however your security team needs.
Use a hosted frontier API for peak quality, or run an open-weight model quantized on your own GPUs for control and cost. Same workflow, same governance — you choose where inference happens.
Inference runs where you want it
One model, four ways to ship it — vendor cloud, your own GPUs, a quantized footprint, or a private VPC endpoint that never leaves your network.
Extended thinking, only where you need it.
Turn on reasoning-grade models for the hard steps — planning, multi-constraint decisions, code, analysis — and keep fast, inexpensive models for routine work. Tune the reasoning effort per node and watch the cost/quality trade-off in real time.
The right model, automatically.
Set a model per node, add fallbacks for resilience, and route by cost, latency, or capability. When a provider rate-limits or a new model ships, switch in one place — no re-platforming.
A multimodal campaign, end to end.
Each node picks its own model and modality. Swap any of them — Whisper for Deepgram, FLUX for Imagen 4, Veo for Sora — without touching the rest of the workflow.
Pick, route, swap.
Pick per node
Assign the best model and modality to each step — frontier, open-weight, or quantized.
Route & fall back
Add fallbacks and route by cost, latency, or capability for resilient runs.
Govern & observe
Every model call is scoped, logged, and traceable — with cost and token usage per step.
Swap without re-platforming
New model ships? Change it in one place. Your workflows keep running.
What teams ask before they commit.
What does it mean that Draft & Goal is model-agnostic?
It means no single AI vendor is baked into the platform. Draft & Goal puts every major provider behind one interface, so you assign the best model to each workflow step — frontier or open-weight — and swap providers without re-platforming. When a new model ships, you change it in one place and your workflows keep running.
Which AI model providers does Draft & Goal support?
Draft & Goal connects models from OpenAI, Anthropic, Google, Meta, Mistral AI, DeepSeek, Alibaba, xAI, and Moonshot AI — including GPT-5, Claude Opus 4.8, Gemini 3 Pro, Llama 4, DeepSeek-R1, and Qwen3. Closed frontier APIs and open-weight models live side by side, so you pick the right one per task.
Can I run open-weight or self-hosted models?
Yes. You can run open-weight models quantized to 4-bit or 8-bit on your own GPUs via vLLM, Ollama, or TGI — in your VPC or on-prem. Draft & Goal also supports private endpoints and bring-your-own API keys, so your security team chooses where inference happens while workflows and governance stay the same.
What modalities can a Draft & Goal workflow handle?
Text, vision, image generation, audio and speech, music, video, and reasoning — composed in a single workflow. One pipeline can transcribe a call with Whisper, reason over it, generate visuals, and render a video, with each node running on the best model for that modality.
How do I balance model cost and quality across a workflow?
Set the model per node: switch on reasoning-grade models for hard steps like planning and analysis, and route fast, inexpensive models for routine work. You can add automatic fallbacks, route by cost, latency, or capability, and every model call is logged with cost and token usage per step.
Show us the workflow.
We'll show you the 10x.
Bring the marketing workflow that eats your week. We'll build it live, with your data and your models, in 30 minutes.