Witbe has developed its own Vision Language Models for automated testing of television and streaming services on physical devices. The models are trained around interface patterns and failure states that general-purpose web models may see rarely, including set-top-box keyboards, launcher grids, older smart TVs and frozen video playback.
The VLMs operate alongside Witbe's agentic framework, which plans and runs scenarios and reacts to the screen. Witbe says specialization reduces latency and cost per test and makes model behavior and upgrade timing more predictable. An on-premises option keeps recorded service video and test data inside the customer's infrastructure.
The system retains a separation between interpretation and proof. AI can navigate and decide what action to take, while deterministic logic validates measured KPIs on the real device and recorded video provides evidence for each run. Operators can also set the AI ratio from zero to 100 percent for individual scenarios.
Rollout has begun and existing Agentic SDK tokens can cover the new models, but detailed infrastructure requirements are available only on request. Witbe says performance on unseen applications is comparable with frontier models, although public benchmark data, model sizes, language coverage and failure rates were not included in the announcement.