Data
The overwhelming majority of illicit personal information trading activities involve cybercrime. Scholars in cybercrime research generally acknowledge that the lack of standardized legal definitions for cybercrime and the absence of effective, reliable official statistics make it challenging to accurately assess the prevalence or incidence of cybercrime globally (McGuire, 2012). Although law enforcement agencies in some countries do collect cybercrime data (e.g., police records and court verdicts), such official data inevitably suffer from underreporting and incomplete documentation (Bossler et al., 2020). This has prompted researchers to turn to alternative data sources for measuring cybercrime, including online transaction forums and cybersecurity firm datasets (Chen et al., 2024).
Illicit personal information trading data are characterized by high sensitivity, fragmentation, and strong anonymity, involving privacy-sensitive information (e.g., biometrics, financial data) traded in batches via dark web markets and encrypted communications, with cryptocurrency used to obscure financial flows. Key challenges in data collection include highly dispersed data(cross-platform, multi-layered resale), technologically adversarial environments (anti-crawling measures, dynamic encryption), and cross-border tracing difficulties (complex judicial coordination, anonymized identity attribution), necessitating integrated multi-source intelligence correlation and legally compliant forensic techniques to overcome barriers (citation needed). Among these data sources, technical datasets—such as crawled data, dynamically encrypted data, intrusion logs, and system logs—are frequently employed as proxy indicators across multiple dimensions of cybercrime and dominate macro-level cybercrime research literature (Loggen et al., 2024; Huang and Deng, 2023).
However, due to the anonymity and virtuality of cyberspace, cybercrime transcends national borders and jurisdictional boundaries, leveraging distributed controlled computers as platforms for illicit personal information trading. Techniques such as proxy servers, anonymous networks (e.g., Tor), virtual private networks (VPNs), and the widespread use of social media apps further complicate the statistical tracking and forensic attribution of such crimes, let alone undetected offenses (Huang and Deng, 2023). While some studies infer illicit personal information trading through related internet black market research, more granular and comprehensive empirical analyses remain scarce.
The data in this study primarily originate from two sources. The first comprises 3,080 telecom fraud case records obtained from the Criminal Investigation Specialized System and Anti-Fraud Platform on the public security internal network of Xiasha District, Zhejiang Province, covering the period from 2013 to 2022. These records include interrogation data (e.g., dialogs between victims and perpetrators, aimed at identifying the types of personal information obtained by criminals) and investigation data (e.g., tracing perpetrators’ methods of contacting victims, such as social media, emails, or phishing websites).
The second category of data comprises 53,100 entries crawled from two Chinese-language dark web platforms: “Dark Web Chinese Marketplace” and “Chang’an Nightless City”. These anonymous platforms, targeting Chinese-speaking users, host extensive illicit trade listings, predominantly involving illicit personal information transactions. These data include fields such as sellers, product titles, product descriptions, US dollar pricing, Bitcoin pricing, transaction volume, and customer reviews.
Analytical descriptive approach
This study employs an analytical descriptive approach (Blei and Lafferty, 2009; Wolniak, 2023) to systematically analyze the operational mechanisms and evolutionary patterns of the personal information black/gray industry chain through the lens of telecom fraud crimes.
Initially, within the descriptive research module, the study investigates the functional role of illicit personal information transactions in criminal supply chains. Through the mining and analysis of case data and dark web transaction data, this research identifies three core functions of illicit personal information transactions in telecom fraud ecosystems.
Subsequently, in the analytical research module, the study explores dynamic evolutionary models of illicit personal information transactions. By applying topic modeling and the temporal sequence analysis method (Chen et al., 2022) to conduct multidimensional modeling of longitudinal crime data spanning 2013–2022, focusing on three evolutionary dimensions: (1) Communication pattern evolution; (2) Transactional modality evolution; (3) Funds flow evolution.
Topic modeling
The implementation workflow of the topic modeling in this study is as follows:
-
Text preprocessing
Tokenization of case transcripts and dark web commodity descriptions (using the Jieba tokenizer).
Stopword filtering (augmented with a criminal investigation terminology stoplist).
Using the NLTK library in Python for lemmatization to reduce noise, lower dimensionality, and enhance topic consistency, thereby improving the performance and efficiency of the model.
-
Hyperparameter tuning
Optimal number of topics (K = 12) determined via perplexity and coherence score evaluation (search range: K = 5–20).
Dirichlet priors: α = 0.1, β = 0.01.
-
Model training
Trained using the Gensim library for 1000 iterations with a convergence threshold of 1e−5.
Extracted top 20 salient terms per topic.
Temporal sequence analysis
Temporal sequence analysis models, mines and forecasts event or state sequences ordered in time, uncovering temporal dependencies and evolutionary laws to portray system dynamics. Temporal sequence analysis models the evolving timing and ordering of cyber-criminal actions to reveal how illicit infrastructures, attack flows and money-laundering chains adapt over time; by mining these temporal patterns across dark-web markets, phishing campaigns and blockchain transactions, defenders can forecast the next tactics or cash-out points and interdict emerging black-gray ecosystems before they scale. This paper’s temporal sequence analysis workflow for the black-and-gray industry proceeds as follows.
-
Phase segmentation
Divided the period 2013–2022 into three intervals: T1 (2013–2015), T2 (2016–2018), T3 (2019–2022).
-
Feature quantification
Tool usage rate: Calculated as the ratio of cases using a specific tool in a given interval to total cases (e.g., telephone/text message usage rate: 91% in T1). and then calculate the technological generational nigration index.
Event Logic Graph approach
An Event Logic Graph (ELG) is a knowledge base of event logic that describes the evolutionary patterns and mechanisms between events (Li et al., 2023; Natthawut and Ryutaro, 2018). The primary purpose of employing ELG in this study is threefold: (1) To model causal-temporal dynamics of criminal activities by formalizing event sequences (e.g., “data theft → dark web transaction → money laundering”) that traditional statistical methods cannot capture; (2) To identify critical leverage points in illicit ecosystems by quantifying relationship weights (e.g., transition probabilities between fraud implementation and fund laundering events); (3) To enable predictive analysis of criminal adaptation patterns through graph-based inference (e.g., simulating impacts of regulatory interventions on transaction evolution).
Structurally, an ELG is a directed cyclic graph where nodes represent events, and directed edges denote logical relationships between events, including temporal succession, causality, conditionality, and hypernymy-hyponymy relations. By integrating the nature, characteristics, and dynamics of illicit personal information trading, event extraction and event relation extraction techniques can construct an ELG for illicit personal information transactions from case data and web-crawled datasets, particularly focusing on three core patterns: “Data transfer patterns,” “Interaction communication patterns,” and “Fund flow patterns,” thereby enabling advanced analysis of their evolutionary mechanisms and lifecycle (Chen et al., 2024).
As a symbolic representation tool for dynamic event evolution, the theoretical framework of ELG is rooted in cognitive linguistics and complex systems theory, using structured knowledge bases to reveal temporal correlations, causal laws, and conditional constraint networks among events (Zhao, 2022). As a novel knowledge representation method in knowledge engineering, ELG emphasizes the construction of event evolution networks with temporality, causality, and logical relevance.
Unlike static knowledge graphs, ELGs capture evolutionary mechanisms through directed edges annotated with transition probabilities. Its technical architecture comprises a node set (E = {e_1, e_2, …, e_n}) and a directed edge set (R = {r_1, r_2, …, r_m}), forming a directed cyclic graph with dynamic evolution features. Nodes represent event entities (e.g., “data theft,” “information resale”), while edge attributes are quantified through binary tuples
Topologically, ELG manifests as a directed cyclic complex network: nodes represent semantically independent event entities (e.g., “data breach,” “illegal transaction”), and directed edges encode multi-level inter-event relationships via predicate logic, including succession (“data collection → cleansing”), causality (“vulnerability exploitation → system intrusion”), conditionality (“dark web platform existence → black-market activity”), and hypernymy (“phishing attacks ⊂ social engineering techniques”) (Pharris and Perez-Mira, 2023; Liao et al., 2022). For illicit personal information trading, ELG construction based on multi-source heterogeneous data (case records, dark web marketplace data) prioritizes three core patterns: (1) Data Transfer Patterns: Chain evolution paths from data theft (e.g., malware implantation) to intermediary resale (e.g., dark web forum transactions) and data cleansing (e.g., anonymization); (2) Interaction Communication Patterns: Criminal group communications via traditional tools, niche social apps, C2 server architectures (e.g., P2P dynamic proxies), encrypted protocols (e.g., Telegram API), and anti-detection tactics (e.g., IP spoofing); (3) Fund Flow Patterns: Cross-border financial pathways involving cash payments, third-party platforms, cryptocurrency mixing (e.g., CoinJoin), and shell company laundering (e.g., offshore account layering) (Shubhdeep and Sukhchandan, 2020; Chen et al., 2024; Xu et al., 2021).
In this study, to establish traceable connections between methodology and findings, the ELG was constructed through four validated stages:
-
Event extraction: From telecom fraud cases (N = 3080) and dark web transactions (N = 53,100), we identified core events using semantic role labeling (e.g., “data theft”, “cryptocurrency payment”) with >92% F1-score in pilot testing.
-
Relationship annotation: Subject matter experts labeled temporal/causal relationships between events (e.g., “phishing → credential capture”) using Cohen’s k = 0.81 inter-rater reliability threshold.
-
Graph assembly: Events became nodes; annotated relationships formed directed edges weighted by transition probability P calculated from case co-occurrence frequency.
-
Validation: The ELG’s precision was verified against 200 ground-truth investigator reports (recall=0.87, precision=0.91).
The methodology flowchart of this study is shown in Fig. 1.

