Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?

Chawla, Divij; Bhutada, Ashita; Anh, Do Duc; Raghunathan, Abhinav; SP, Vinod; Guo, Cathy; Liew, Dar Win; Gupta, Prannaya; Bhardwaj, Rishabh; Bhardwaj, Rajat; Poria, Soujanya

Computer Science > Computation and Language

arXiv:2505.18953 (cs)

[Submitted on 25 May 2025]

Title:Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?

Authors:Divij Chawla, Ashita Bhutada, Do Duc Anh, Abhinav Raghunathan, Vinod SP, Cathy Guo, Dar Win Liew, Prannaya Gupta, Rishabh Bhardwaj, Rajat Bhardwaj, Soujanya Poria

View PDF HTML (experimental)

Abstract:We evaluate the credibility of leading AI models in assessing investment risk appetite. Our analysis spans proprietary (GPT-4, Claude 3.7, Gemini 1.5) and open-weight models (LLaMA 3.1/3.3, DeepSeek-V3, Mistral-small), using 1,720 user profiles constructed with 16 risk-relevant features across 10 countries and both genders. We observe significant variance across models in score distributions and demographic sensitivity. For example, GPT-4o assigns higher risk scores to Nigerian and Indonesian profiles, while LLaMA and DeepSeek show opposite gender tendencies in risk classification. While some models (e.g., GPT-4o, LLaMA 3.1) align closely with expected scores in low- and mid-risk ranges, none maintain consistent performance across regions and demographics. Our findings highlight the need for rigorous, standardized evaluations of AI systems in regulated financial contexts to prevent bias, opacity, and inconsistency in real-world deployment.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2505.18953 [cs.CL]
	(or arXiv:2505.18953v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2505.18953

Submission history

From: Rishabh Bhardwaj [view email]
[v1] Sun, 25 May 2025 02:56:19 UTC (478 KB)

Computer Science > Computation and Language

Title:Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators