Introducing Velvet Flash 0.1 (For Velvet-1): Smaller, Specialized, Better
We’ve launched Velvet Flash 0.1, our own 4B model built specifically for crypto skills. Despite being much smaller, it ranked #1 in our benchmark and outperformed models up to 70B on crypto tasks
Today, we are introducing Velvet Flash 0.1, our in-house, task-specialized 4-billion-parameter model built for safe and reliable crypto operations.
We created Velvet Flash to handle the parts of crypto interaction where general-purpose models often struggle. It needs to understand user intent, select the correct platform command, extract parameters accurately, recognize risky requests, and confirm before any action involving funds. A model must be precise, consistent, and safe.
On our crypto-skills benchmark, Velvet Flash ranked first overall with a score of 50. Qwen3.5-27B scored 45, DeepSeek-V4 scored 40, Llama-3.3-70B scored 32, Gemma-4-31B scored 31, and Mistral-24B scored 20. Velvet Flash achieved the highest score despite having only 4 billion parameters, making it between 6 and 17 times smaller than several of the models it outperformed.
The benchmark evaluates how well a model behaves as a safe crypto operator. Each model receives a platform skill document together with a natural-language request. It must choose the correct command, extract parameters such as amounts, tokens, chains, and recipient addresses, confirm before any fund-moving action, and warn about scams, wrong chains, or risky requests.
The clearest evidence of what training achieved is the difference between Velvet Flash and its untrained base model. The base model scored 23, while Velvet Flash scored 50. This means task-specific training more than doubled the overall score.
The model was also evaluated across five capability dimensions: safety, coverage, robustness, routing, and clarity. Safety improved from 28 to 61, robustness from 26 to 53, clarity and UX from 22 to 52, routing from 30 to 47, and coverage from 12 to 34.
The safety result matters most to us. In a financial product, a model must know when to continue, when to ask for clarification, when to warn the user, and when not to act. It must parse amounts correctly, protect sensitive information, identify suspicious requests, and avoid moving funds without confirmation. These behaviors are not secondary product features.
Some larger models performed better on individual platforms. Qwen3.5-27B and Llama-3.3-70B scored higher on Minara, while Qwen also performed better on Binance Spot.
Model quality is only part of the story. A 4B model is easier and less expensive to serve than a 27B, 31B, or 70B model. It can offer lower latency, require less infrastructure, and give us greater control over deployment and iteration. Velvet Flash is served on our own infrastructure, which allows us to optimize the full system around the needs of the product rather than relying on a general model as an external black box.
The benchmark uses a six-platform core set covering Minara, Binance Spot, OKX DEX, Uniswap, GMX, and MetaMask. The broader Velvet evaluation spans 31 platforms across centralized exchanges, wallets, DEXs, DeFi products, and trading tools. All models received identical inputs, identical held-out scenarios, and the same independent judge. Responses were scored across five weighted dimensions, with an additional safety gate for failures such as moving funds without confirmation or misparsing an amount.
Velvet Flash 0.1 is the first version of this model. We will continue expanding platform coverage, improving difficult edge cases, strengthening safety behavior, and refining performance on complex multi-step workflows.
The broader goal is to build an internal capability for creating specialized models that are safer, faster, cheaper, and better aligned with the products they serve.








