Project
DEA Evolution and Bibliometric Pipeline
An end-to-end pipeline for cleaning, classifying, and mapping 54K Scopus records through scholar metrics, collaboration networks, and main-path analysis.
Python · Scopus · Bibliometrics · Network Analysis · DeepSeek
Product Snapshot
- Role
- Data research and analysis engineering
- Users
- Research teams studying the evolution, applications, and knowledge structure of DEA.
- Stage
- Data cleaning, classification, metrics, network analysis, and report assets completed in 2025.09-2026.03.
- Focus
- Large-scale bibliographic cleaning, LLM-assisted classification, scholar metrics, networks, and SPC main paths.
- Validation
- Reduced 53,954 raw Scopus records to about 33,690 application papers and produced a reusable analysis pipeline.
- Public Proof
- Raw and processed datasets, Python scripts, figures, and research outputs.
Next Step
Research goal
The project asked how DEA moved from theory into application domains, which authors and institutions formed key networks, and how the field’s main paths changed over time.
Method
I cleaned and normalized 53,954 Scopus records, producing about 33,690 application papers. The pipeline used DeepSeek for assisted domain classification, calculated h/g-index measures, built collaboration and similarity networks, and extracted SPC main paths.
Result and boundary
The work produced reusable Python analysis, structured datasets, and visual outputs. LLM classification expanded coverage but did not replace sampled checks, so the findings describe distributions and network structure rather than treating automated labels as ground truth.
Configure the public Giscus environment variables to open discussion on the public site.