Witbe has detailed an agentic testing and monitoring stack that runs on real smart TVs, set-top boxes, mobile phones and streaming sticks. Its Agentic SDK mixes natural-language, AI-driven steps with deterministic checks for startup time, buffering and mean-opinion-score video quality.
A peer-reviewed IBC paper covers eight months of production use. According to Witbe, one engineer grew from 78 to 316 maintained tests and completed 280,626 executions while agents adapted to interface changes. The platform also tests vertical video, virtual advertisements, stay-live advertising and captions as rendered on screen rather than relying only on back-end signals.
The approach could reduce brittle image-matching scripts while preserving numerical checks where repeatability matters. The results are vendor-reported and tied to a specific deployment; broader reproducibility, model cost, failure rates and the boundaries of autonomous adaptation remain open.