United Overseas Bank (UOB)
Singapore · Banking · 2025
In the AI Verify Foundation's Global AI Assurance Pilot (February to May 2025), PwC tested UOB's internal retrieval augmented generation chatbot, which runs in production for selected staff on Meta Llama 3.1 and answers operational and domain questions from public company documents. The risk assessment focused on model risks. PwC combined rule based scoring for binary and multiple choice questions, embedding similarity for consistency across repeated runs, and LLM based checks of reasoning answers: an LLM split each answer into clauses, an LLM as a judge compared each clause with retrieved passages of the source document to flag contradictions (a clause with no supporting passage counted as a hallucination), and a judge listed the parts of each question left unanswered. Because the production infrastructure was shared with other use cases, outputs were generated manually in a sandbox; because of confidentiality, PwC used its own prompts and ground truths for ten companies, which UOB reviewed. The case study therefore treats the results as a proxy for the production tool and publishes none of them.
No outcome disclosed.