Language-learning AI tools are easy to demo well and hard to evaluate properly — a slick chat interface can hide whether it's actually teaching anything durable. A few criteria matter more than how polished the product looks.
Pedagogical grounding. Does the tool follow any recognised approach to language acquisition, or is it just a general chatbot wrapped in a language-learning skin? The difference shows up in whether it reinforces, corrects and spaces out practice — or just answers whatever's typed.
Accuracy in the target language pair. Large language models are noticeably stronger in some languages than others. A tool that performs well in English-Spanish practice isn't guaranteed to perform equally well in less commonly taught languages — this needs testing, not assuming.
Correction, not just conversation. Free-flowing chat is engaging but doesn't automatically teach. Look for tools that explicitly flag and explain errors, rather than quietly understanding broken grammar and moving on.
Teacher visibility. In an institutional setting, a tool that gives teachers no insight into what students are actually practising or getting wrong is a black box — useful for engagement, harder to justify pedagogically.
Realistic accessibility. Voice-based practice tools in particular need to be checked against students who rely on screen readers or have different accents than the tool was trained on.
The honest test: would this tool's value survive being used for a full term, not just a five-minute demo?