Find your way to be trained or even get certified in TMAP.
Start typing keywords to search the site. Press enter to submit.
Testing AI with AI is a special case of using artificial intelligence. It is about using AI to test systems that are based on (generative or other) artificial intelligence. This approach is commonly known as “AI as a judge,” or more specifically, in generative AI contexts, “LLM as a judge.”
This may be a very appealing and sometimes feasible possibility. However valid this option may seem, before deciding to do so, the people involved must very carefully weigh the fact that they will test a system of which they don’t exactly know what it does, by using another system of which they don’t exactly know what it does. So, all in all this is piling up uncertainties. On the other hand, you may argue that in our modern systems of systems, we have long ago become used to trusting systems we don’t fully understand.
The key is in mitigating the risks. This is done by keeping an expert in the lead. An expert that knows what the system under test is supposed to do, and who understands how artificial intelligence can be applied to test such system, and who is able to recognize whether the quality level that is reported actually reflects the quality level that a human will perceive.The final judgment about confidence that the pursued business value will be achieved still has to be made by a human being, because the accountability for business processes can’t be delegated to a machine, no matter how intelligent it may seem.
One of the key challenges human experts face when determining the expected outcome of tests is the so-called oracle problem. This problem arises when testers must judge whether a system’s behavior is correct in situations where the task itself is broadly defined, which is common for GenAI systems. Unlike traditional deterministic software, GenAI outputs are probabilistic and context-dependent, making it difficult to define a single “correct” expected outcome. Vague or incomplete requirements make the oracle problem even worse, as they leave significant room for interpretation. To mitigate this, quality engineers can involve domain experts who contribute deep contextual knowledge to establish more meaningful acceptance criteria. In addition, statistical testing techniques can be applied to assess output quality and consistency across large sample sets rather than individual results. Finally, the inherent creativity of a GenAI model can be constrained by tuning technical parameters such as the temperature, allowing teams to trade off variability for predictability when testability and reliability are critical.
One of the risks associated with using AI is that people don’t apply critical thinking when assessing the results of an AI based solution. People are quickly misled by the nice-looking outputs or may even lack the knowledge needed (and may also lack curiosity) to judge the validity of output. This is known as the “automation bias” of people which means that if a result is produced by an automated system, people tend to trust it highly.
One of the results of testing an AI-based solution may be to create a usage guideline that describes in which situations the solution can be used and in which situations it is not wise (or sometimes even illegal) to use the solution.This can help raise the awareness of the people involved about the risks an AI-based solution brings.
Read more about ‘Critical Thinking‘.
AmpQE overview