Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问许可证需确认审计通过

stata-data-cleaningstata 数据清理

Agent Skill

用于辅助数据整理、表格处理、CSV/Excel 分析、指标计算和图表准备。它适合让 Agent 清洗字段、汇总数据、发现异常、生成统计口径或把分析结果转成可读说明。使用时需要确认数据来源、字段含义和时间范围,避免把样本数据当全量事实;涉及敏感数据、导出文件或批量写回时,应先确认权限和脱敏边界。

总安装

3,469

周安装

149

GitHub Stars

375

下载量

1,216
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:stata-data-cleaning(stata 数据清理)
来源仓库:https://github.com/meleantonio/awesome-econ-ai-stuff
仓库路径:skills/stata-data-cleaning
安装命令:
npx skills add https://github.com/meleantonio/awesome-econ-ai-stuff --skill stata-data-cleaning
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/meleantonio/awesome-econ-ai-stuff --skill stata-data-cleaning

简介

stata-data-cleaning 用于辅助数据整理、CSV/Excel 分析和指标计算,适合在 Codex、Claude、Cursor、Gemini CLI 中清洗字段和发现异常。

  • 适用于数据预处理、统计口径生成和图表准备等场景,如经济数据分析、数据报告。
  • 通过 npx skills add 命令从 GitHub 仓库安装,需结合数据来源和字段含义使用。
  • 涉及敏感数据或批量写回时应先确认权限和脱敏边界,避免数据泄露。
  • 使用前建议核对时间范围和样本数据,防止将局部结果当作全量事实。

SKILL.md

Stata Data Cleaning

Purpose

This skill helps economists clean, transform, and prepare datasets for analysis in Stata. It emphasizes reproducibility, proper documentation, and handling common data quality issues found in economic research.

When to Use

  • Cleaning raw survey or administrative data
  • Merging multiple data sources
  • Handling missing values, duplicates, and outliers
  • Creating analysis-ready panel datasets
  • Documenting data transformations for replication

Instructions

Step 1: Understand the Data

Before generating code, ask the user:

  1. What is the data source? (survey, administrative, API, etc.)
  2. What is the unit of observation?
  3. What are the key variables needed for analysis?
  4. Are there known data quality issues to address?

Step 2: Generate Cleaning Pipeline

Create a Stata do-file that:

  1. Has a clear header with project info and date
  2. Sets up the environment (clear all, set memory, log)
  3. Loads and inspects raw data
  4. Documents each transformation with comments
  5. Creates a codebook for the final dataset

Step 3: Follow Best Practices

  • Use assert statements to verify data integrity
  • Create labeled variables with label variable
  • Use value labels for categorical variables
  • Generate a log file for reproducibility
  • Save intermediate files when appropriate

Example Output

/*==============================================================================
    Project:    Economic Analysis Data Cleaning
    Author:     [Your Name]
    Date:       [Date]
    Purpose:    Clean raw survey data for regression analysis
    Input:      raw_survey_data.dta
    Output:     cleaned_analysis_data.dta
==============================================================================*/

* ============================================
* 1. SETUP
* ============================================

clear all
set more off
cap log close
log using "logs/data_cleaning_`c(current_date)'.log", replace

* Set working directory
cd "/path/to/project"

* Define globals for paths
global raw_data "data/raw"
global clean_data "data/clean"
global output "output"

* ============================================
* 2. LOAD AND INSPECT RAW DATA
* ============================================

use "${raw_data}/raw_survey_data.dta", clear

* Basic inspection
describe
summarize
codebook, compact

* Check for duplicates
duplicates report id_var
duplicates list id_var if _dup > 0

* ============================================
* 3. VARIABLE CLEANING
* ============================================

* --- Rename variables for clarity ---
rename q1 age
rename q2 income_reported
rename q3 education_level

* --- Clean numeric variables ---
* Replace missing value codes with .
mvdecode age income_reported, mv(-99 -88 -77)

* Cap outliers at 99th percentile
qui sum income_reported, detail
replace income_reported = r(p99) if income_reported > r(p99) & !mi(income_reported)

* --- Clean string variables ---
* Standardize state names
replace state = upper(trim(state))
replace state = "NEW YORK" if inlist(state, "NY", "N.Y.", "N Y")

* --- Create categorical variables ---
gen education_cat = .
replace education_cat = 1 if education_level < 12
replace education_cat = 2 if education_level == 12
replace education_cat = 3 if education_level > 12 & education_level <= 16
replace education_cat = 4 if education_level > 16 & !mi(education_level)

label define edu_lbl 1 "Less than HS" 2 "High School" 3 "College" 4 "Graduate"
label values education_cat edu_lbl

* ============================================
* 4. HANDLE MISSING DATA
* ============================================

* Create missing indicator variables
gen mi_income = mi(income_reported)

* Document missingness
tab mi_income

* Count complete cases
egen complete_case = rownonmiss(age income_reported education_cat)
tab complete_case

* ============================================
* 5. CREATE DERIVED VARIABLES
* ============================================

* Age groups
gen age_group = .
replace age_group = 1 if age >= 18 & age < 30
replace age_group = 2 if age >= 30 & age < 50
replace age_group = 3 if age >= 50 & age < 65
replace age_group = 4 if age >= 65 & !mi(age)

label define age_lbl 1 "18-29" 2 "30-49" 3 "50-64" 4 "65+"
label values age_group age_lbl

* Log income
gen log_income = ln(income_reported + 1)

* ============================================
* 6. DATA VALIDATION
* ============================================

* Assert expected ranges
assert age >= 18 & age <= 120 if !mi(age)
assert income_reported >= 0 if !mi(income_reported)

* Check variable types
assert !mi(id_var)
isid id_var  // Verify unique identifier

* ============================================
* 7. LABEL VARIABLES
* ============================================

label variable age "Age in years"
label variable income_reported "Annual income (USD)"
label variable education_cat "Education category"
label variable log_income "Log of annual income"
label variable mi_income "Missing income indicator"

* ============================================
* 8. FINAL CHECKS AND SAVE
* ============================================

* Keep relevant variables
keep id_var age age_group income_reported log_income ///
     education_cat mi_income state year

* Order variables logically
order id_var year state age age_group income_reported ///
      log_income education_cat mi_income

* Compress to minimize file size
compress

* Save cleaned data
save "${clean_data}/cleaned_analysis_data.dta", replace

* Create codebook
codebook, compact

* Close log
log close

* ============================================
* END OF FILE
* ============================================

Requirements

Software

  • Stata 15+ (some commands require newer versions)

Recommended User-Written Commands

ssc install unique     // For unique value checking
ssc install mdesc      // For missing data patterns
ssc install labutil    // For label manipulation

Best Practices

  1. Always start with clear all to ensure clean environment
  2. Use log files to document all transformations
  3. Comment extensively - explain WHY, not just WHAT
  4. Use assert statements to catch data errors early
  5. Create a data dictionary alongside your cleaned data
  6. Version your do-files and datasets

Common Pitfalls

  • ❌ Not checking for duplicates before merging
  • ❌ Forgetting to handle missing value codes (-99, -88, etc.)
  • ❌ Not labeling variables and values
  • ❌ Overwriting raw data files
  • ❌ Not documenting data transformations

References

Changelog

v1.0.0

  • Initial release with comprehensive cleaning template

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

37.96%
按下载量换算462

Claude

29.34%
按下载量换算357

Cursor

19.4%
按下载量换算236

Gemini CLI

8.91%
按下载量换算108

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills