Leveraging Artificial Intelligence and Machine Learning for Real-Time Fraud Detection in E-Commerce Transactions
Sachin Bagoria1*, Dr. Kavita2
  1. Research Scholar, Shri Khushal Das University, Hanumangarh, Rajasthan, India
radheykrishnalalita@gmail.com
2 Professor, Shri Khushal Das University, Hanumangarh, Rajasthan, India
Abstract: With more and more people making purchases online, fraud detection has become an important issue for online marketplaces due to the explosion of e-commerce. However, traditional approaches often are inadequate for catching more sophisticated and emerging fraudulent activities in real-time. The aim of this research is to investigate the effectiveness of ML and AI techniques in realtime detection of online shopping fraud. The number of 50,000 records were selected from marketplaces using stratified sampling to ensure representative sampling of classes. Data preparation involved dealing with missing data, removing duplicate data, address outliers, scaling features, and encoding categorical data for analysis. To overcome the problem of class imbalance, the use of SMOTENC, SMOTENC + ENN, & SMOTENC + Tomek Links approaches were performed.
Numerous ML classifiers, such as Random Forest & Stochastic Gradient, were tested. The model was assessed for several parameters such as recall, accuracy, precision, F1 score, and AUC-ROC. Random Forest (RF) out performed all the other classifiers in both balanced and unbalanced datasets and Stochastic Gradient (SG) performed the next best. The most important factors that go into fraud detection judgements were determined via SHAP analysis. The study highlights potential opportunities for trust in e-commerce platforms, risk mitigations in terms of financial losses, and enhanced transaction security through AI-driven fraud detection.
Keywords: Artificial Intelligence, Machine Learning, Fraud Detection, E-Commerce Transactions, Random Forest, SMOTENC, SHAP Analysis, Class Imbalance, Cybersecurity, Predictive Analytics.

    INTRODUCTION

In recent ten years, e-commerce has emerged as one of the fastest-growing segments in the global economy, growing from around 0.5% of GDP in the world to more than 1.5% [1]. There are advantages that companies and customers can offer to online shopping and digital payment systems & electronic transactions and those possibilities have been expanding rapidly in recent years. However, this growth has been accompanied by the emergence of cybercrime and fraudulent activities. Due to the increasing danger of online fraud, the projected worldwide cost of cybercrime jumped from $445 billion in 2014 to over $600 billion in 2017.
Examples of some of the many forms of malicious activity encompassed by the term 'ecommerce fraud' include: fraudulent listing creation, fraudulent reviews, account takeover attacks, fraudulent payments and fraudulent accounts [3, 4]. This has caused businesses to deal with massive monetary losses and online marketplaces lose credibility and consumer confidence. As the number of transactions has increased at an exponential rate, the traditional rule-based fraud detection systems which fail to detect complex and evolving fraud scenarios have also become inadequate.
The use of AI and ML in detecting, preventing and mitigating fraudulent activities on online marketplaces has increased by a huge margin. Fraud detection systems are widely adopted by large companies like Microsoft [5], LinkedIn [6] and eBay [7] that use machine learning to make their systems more efficient. These systems are able to quickly detect any fraudulent actions by analysing massive amounts of behavioural and transactional data, finding unusual patterns, and drawing conclusions.
Problems with fraud datasets' extreme imbalance and the ever-increasing complexity of fraud schemes persist despite substantial progress in the field of fraud detection. It is challenging for machine learning algorithms to correctly detect minority-class occurrences since fraudulent transactions usually only account for a tiny proportion of overall transactions. Therefore, using good data preparation, class balancing tactics, feature engineering, & model interpretability methodologies is crucial for developing strong fraud detection systems.
Within this framework, the current research delves into the use of AI and ML for the purpose of detecting online transaction fraud in real-time. It features a comprehensive experimental setup, data preparation, stratified sampling, stratified majority minority sampling (SMOTENC) techniques to handle class imbalance, and evaluation of multiple machine learning classifiers. In addition, the most important characteristics that go into fraud detection judgements are identified using SHAP analysis. The aim of this study is to propose a classification algorithm and rebalancing procedures comparison that can be useful in the design of a practical and understandable solution to improve the performance of e-commerce fraud detection.
This study's findings are expected to contribute to the creation of a reliable AI-powered fraud detection system, which will help enable safe online transactions, minimise financial losses and boost consumer confidence in online shopping platforms.
Sachin Bagoria, Dr. Kavita

2. OBJECTIVES

In order to create and assess ML and AI models for e-commerce fraud detection in real-time, this study uses a quantitative & experimental research approach. Examining numerical and categorical transaction variables and evaluating the performance of different machine learning algorithms in various data-balancing procedures is the main emphasis of the study, which aims to uncover fraudulent actions [8].

3.1 Data Collection

Listing information and transaction data from an online marketplace are included in the data collection. For fair and efficient distribution the selected sample of 50,000 records were randomly selected from the overall data set using a stratified sampling technique. Using stratified sampling, we were able to guarantee that our sample was representative of all types of transactions and fraud.

3.2 Data Preprocessing

A comprehensive pretreatment pipeline was developed to prepare the data for the analysis under machine learning. Some of the preprocessing techniques employed include: