R Programming Anti-Slop Skill
When to Use This Skill
Use r-anti-slop when:
- ✓ Writing new R code for data analysis or packages
- ✓ Reviewing AI-generated R code before committing
- ✓ Refactoring existing code for production quality
- ✓ Preparing R package for CRAN submission
- ✓ Teaching or enforcing R code standards
- ✓ Cleaning up generic variable names and patterns
Do NOT use when:
- Writing quick exploratory one-offs (though standards still help)
- Working with legacy code that cannot be changed
- Following different established style guides (e.g., Bioconductor)
Quick Example
Before (AI Slop):
# Load the library
library(dplyr)
# Read the data
df <- read.csv("data.csv")
# Filter the data
result <- df %>% filter(x > 0)
After (Anti-Slop):
customer_data <- readr::read_csv("data/customers.csv")
active_customers <- customer_data |>
dplyr::filter(status == "active", revenue > 0)
return(active_customers)
What changed:
- ✓ Descriptive names (
customer_data not df)
- ✓ Namespace qualification (
dplyr::, readr::)
- ✓ Native pipe (
|> not %>%)
- ✓ No obvious comments
- ✓ Explicit return
When to Use What
| If you need to... |
Do this |
Details |
| Name variables |
Use snake_case, no df/data/result |
reference/naming.md |
| Call tidyverse functions |
Always use :: (e.g., dplyr::filter()) |
reference/tidyverse.md |
| Return from function |
Always explicit return() statement |
reference/naming.md |
| Write pipe chains |
Use |>, break at 8+ operations |
reference/tidyverse.md |
| Document functions |
Specific @param, @return, no circular text |
reference/documentation.md |
| Handle missing data |
Explicit strategy + report data loss |
reference/statistical-rigor.md |
| Validate data |
Check assumptions with stopifnot() |
reference/statistical-rigor.md |
| Format code |
Use styler::style_file() |
reference/tidyverse.md |
| Check code quality |
Use lintr::lint() |
reference/tidyverse.md |
Core Workflow
5-Step Quality Check
Namespace qualification - All external functions use ::
# Good
dplyr::filter(data, x > 0)
# Bad
filter(data, x > 0)
Explicit returns - Every function has return()
# Good
my_function <- function(x) {
result <- x + 1
return(result)
}
# Bad
my_function <- function(x) {
x + 1
}
Naming conventions - All objects use snake_case
# Good
customer_lifetime_value <- calculate_clv(data)
# Bad
df <- calculate_clv(data)
customerLifetimeValue <- calculate_clv(data)
Documentation quality - No generic descriptions
# Good
#' @param deaths Data frame with `age_group` and `count` columns
# Bad
#' @param data The data
Code formatting - Run styler and lintr
styler::style_file("script.R")
lintr::lint("script.R")
Quick Reference Checklist
Before committing R code, verify:
Common Workflows
Workflow 1: Clean Up AI-Generated R Script
Context: AI generated an analysis script with generic patterns.
Steps:
Run detection script
Rscript toolkit/scripts/detect_slop.R analysis.R --verbose
Fix high-priority issues first
# Replace df, data, result with descriptive names
# Before
df <- readr::read_csv("data.csv")
result <- df %>% filter(x > 0)
# After
customer_data <- readr::read_csv("data/customers.csv")
active_customers <- customer_data |> dplyr::filter(status == "active")
Add namespace qualification
# Before
data %>% filter(x > 0) %>% summarize(mean(y))
# After
data |>
dplyr::filter(x > 0) |>
dplyr::summarize(mean_y = mean(y))
Add explicit returns
# Before
calculate_rate <- function(numerator, denominator) {
numerator / denominator
}
# After
calculate_rate <- function(numerator, denominator) {
rate <- numerator / denominator
return(rate)
}
Break long pipes
# Before (12 operations in one chain)
result <- data |>
filter(...) |> mutate(...) |> group_by(...) |>
summarize(...) |> arrange(...) |> [7 more ops]
# After
clean_data <- data |>
dplyr::filter(!is.na(value)) |>
dplyr::mutate(category = categorize(value))
summary_stats <- clean_data |>
dplyr::group_by(category) |>
dplyr::summarize(mean_val = mean(value))
Format and validate
styler::style_file("analysis.R")
lintr::lint("analysis.R")
Expected outcome: Score drops from 60+ to <20
Workflow 2: Fix Generic Package Documentation
Context: R package has generic roxygen documentation.
Steps:
Identify generic patterns
# Bad
#' Process Data
#'
#' @description This function processes the data.
#' @param data The data.
#' @return The result.
Make description specific
# Good
#' Calculate age-adjusted mortality rates
#'
#' Computes mortality rates per 100,000 population, standardized to the
#' 2000 US Census age distribution using direct standardization.
Describe parameter structure
# Good
#' @param deaths Data frame with columns `age_group` and `count`.
#' @param population Data frame with columns `age_group` and `pop_size`.
Specify return value
# Good
#' @return A tibble with columns:
#' \describe{
#' \item{county}{County FIPS code}
#' \item{rate}{Age-adjusted rate per 100,000}
#' \item{se}{Standard error of the rate}
#' }
Add realistic examples
# Good
#' @examples
#' counties <- data.frame(
#' county = c("A", "B"),
#' deaths = c(150, 200),
#' population = c(50000, 80000)
#' )
#'
#' adjust_rates(counties, rate_per = 100000)
#' #> # A tibble: 2 x 3
#' #> county rate se
#' #> 1 A 312. 25.4
#' #> 2 B 258. 18.2
Expected outcome: Documentation that teaches, not restates
Workflow 3: Prepare Package for CRAN
Context: Final checks before CRAN submission.
Steps:
Run all quality checks
# Standard checks
devtools::check()
# Anti-slop checks
lapply(list.files("R", full.names = TRUE), function(f) {
system(paste("Rscript toolkit/scripts/detect_slop.R", f))
})
Fix documentation
- Check all
@param descriptions are specific
- Verify
@examples run and are realistic
- Ensure
@return describes structure
Validate code quality
# Format all files
styler::style_dir("R/")
# Check lints
lintr::lint_package()
Check CRAN-specific requirements
- Use external/posit-skills/r-lib/cran-extrachecks skill
- Validate URLs in DESCRIPTION and documentation
- Check examples run in < 5 seconds
Expected outcome: Clean R CMD check with no slop patterns
Mandatory Rules Summary
1. Namespace Qualification
ALWAYS use :: for external packages
Exceptions (don't need ::):
- Base R:
mean(), sum(), log(), etc.
- stats:
lm(), glm(), t.test(), etc.
- utils:
head(), tail(), str(), etc.
2. Explicit Returns
ALWAYS use return() - never implicit
3. Naming: snake_case
All objects use snake_case
- Variables:
customer_data not customerData or df
- Functions:
calculate_rate not calculateRate
- Arguments:
input_data not inputData
4. Native Pipe
Prefer |> over %>% (unless R < 4.1)
5. No Generic Names
Never use: df, data, result, temp, x, n (except standard math notation)
Tidyverse Philosophy
Follow Tidyverse Style Guide as primary reference:
- Design for humans - Code should be readable and intuitive
- Reuse existing data structures - Work with tibbles and data frames
- Compose simple functions with pipes - Build complexity through composition
- Embrace functional programming - Functions are first-class objects
See reference/tidyverse.md for complete tidyverse conventions.
Resources & Advanced Topics
Reference Files
- reference/naming.md - Complete naming conventions and forbidden patterns
- reference/tidyverse.md - Pipe conventions, formatting, ggplot2 standards
- reference/documentation.md - Roxygen2, vignettes, README quality
- reference/statistical-rigor.md - Validation, uncertainty, reproducibility
- reference/forbidden-patterns.md - Complete antipattern catalog
Related Skills
- external/posit-skills/r-lib/cli - Error message formatting with cli package
- external/posit-skills/r-lib/testing - Test structure and best practices
- external/posit-skills/r-lib/cran-extrachecks - CRAN submission requirements
- text/anti-slop - For cleaning prose in documentation
Tools
styler::style_file() - Auto-format code
lintr::lint() - Check code quality
Rscript toolkit/scripts/detect_slop.R - Detect AI patterns
Integration with Posit Skills
This skill focuses on code quality and avoiding generic patterns.
Use together with Posit skills for complete coverage:
| Task |
Use This Skill |
+ Posit Skill |
| Write error messages |
r/anti-slop (quality) |
+ r-lib/cli (structure) |
| Write tests |
r/anti-slop (code quality) |
+ r-lib/testing (test patterns) |
| Prepare for CRAN |
r/anti-slop (no slop) |
+ r-lib/cran-extrachecks (requirements) |
| Document lifecycle |
r/anti-slop (doc quality) |
+ r-lib/lifecycle (deprecation) |
1---2name: r-anti-slop-23description: Enforce production-quality R code standards. Prevents generic AI patterns through namespace qualification, explicit returns, and tidyverse conventions. Use when writing or reviewing R code for data analysis or packages.4---56# R Programming Anti-Slop Skill78## When to Use This Skill910Use r-anti-slop when:11- ✓ Writing new R code for data analysis or packages12- ✓ Reviewing AI-generated R code before committing13- ✓ Refactoring existing code for production quality14- ✓ Preparing R package for CRAN submission15- ✓ Teaching or enforcing R code standards16- ✓ Cleaning up generic variable names and patterns1718Do NOT use when:19- Writing quick exploratory one-offs (though standards still help)20- Working with legacy code that cannot be changed21- Following different established style guides (e.g., Bioconductor)2223## Quick Example2425**Before (AI Slop)**:26```r27# Load the library28library(dplyr)2930# Read the data31df <- read.csv("data.csv")3233# Filter the data34result <- df %>% filter(x > 0)35```3637**After (Anti-Slop)**:38```r39customer_data <- readr::read_csv("data/customers.csv")4041active_customers <- customer_data |>42 dplyr::filter(status == "active", revenue > 0)4344return(active_customers)45```4647**What changed**:48- ✓ Descriptive names (`customer_data` not `df`)49- ✓ Namespace qualification (`dplyr::`, `readr::`)50- ✓ Native pipe (`|>` not `%>%`)51- ✓ No obvious comments52- ✓ Explicit return5354## When to Use What5556| If you need to... | Do this | Details |57|-------------------|---------|---------|58| Name variables | Use `snake_case`, no `df`/`data`/`result` | reference/naming.md |59| Call tidyverse functions | Always use `::` (e.g., `dplyr::filter()`) | reference/tidyverse.md |60| Return from function | Always explicit `return()` statement | reference/naming.md |61| Write pipe chains | Use `\|>`, break at 8+ operations | reference/tidyverse.md |62| Document functions | Specific `@param`, `@return`, no circular text | reference/documentation.md |63| Handle missing data | Explicit strategy + report data loss | reference/statistical-rigor.md |64| Validate data | Check assumptions with `stopifnot()` | reference/statistical-rigor.md |65| Format code | Use `styler::style_file()` | reference/tidyverse.md |66| Check code quality | Use `lintr::lint()` | reference/tidyverse.md |6768## Core Workflow6970### 5-Step Quality Check71721. **Namespace qualification** - All external functions use `::`73 ```r74 # Good75 dplyr::filter(data, x > 0)76 # Bad77 filter(data, x > 0)78 ```79802. **Explicit returns** - Every function has `return()`81 ```r82 # Good83 my_function <- function(x) {84 result <- x + 185 return(result)86 }87 # Bad88 my_function <- function(x) {89 x + 190 }91 ```92933. **Naming conventions** - All objects use `snake_case`94 ```r95 # Good96 customer_lifetime_value <- calculate_clv(data)97 # Bad98 df <- calculate_clv(data)99 customerLifetimeValue <- calculate_clv(data)100 ```1011024. **Documentation quality** - No generic descriptions103 ```r104 # Good105 #' @param deaths Data frame with `age_group` and `count` columns106 # Bad107 #' @param data The data108 ```1091105. **Code formatting** - Run styler and lintr111 ```r112 styler::style_file("script.R")113 lintr::lint("script.R")114 ```115116## Quick Reference Checklist117118Before committing R code, verify:119120- [ ] All external functions qualified with `::`121- [ ] All functions have explicit `return()`122- [ ] All objects use `snake_case`123- [ ] No generic names (`df`, `data`, `result`, `temp`)124- [ ] Pipes (`|>`) have space before, end lines125- [ ] Long pipelines (>8 ops) broken into named steps126- [ ] Complex operations have WHY comments127- [ ] Data validated after transformations128- [ ] Seeds set before random operations129- [ ] Uncertainty reported (SE, CI) for statistical models130- [ ] No `attach()` calls131- [ ] No right-hand assignment (`->`)132- [ ] Roxygen documentation is specific133- [ ] Examples are realistic and run134135## Common Workflows136137### Workflow 1: Clean Up AI-Generated R Script138139**Context**: AI generated an analysis script with generic patterns.140141**Steps**:1421431. **Run detection script**144 ```bash145 Rscript toolkit/scripts/detect_slop.R analysis.R --verbose146 ```1471482. **Fix high-priority issues first**149 ```r150 # Replace df, data, result with descriptive names151 # Before152 df <- readr::read_csv("data.csv")153 result <- df %>% filter(x > 0)154155 # After156 customer_data <- readr::read_csv("data/customers.csv")157 active_customers <- customer_data |> dplyr::filter(status == "active")158 ```1591603. **Add namespace qualification**161 ```r162 # Before163 data %>% filter(x > 0) %>% summarize(mean(y))164165 # After166 data |>167 dplyr::filter(x > 0) |>168 dplyr::summarize(mean_y = mean(y))169 ```1701714. **Add explicit returns**172 ```r173 # Before174 calculate_rate <- function(numerator, denominator) {175 numerator / denominator176 }177178 # After179 calculate_rate <- function(numerator, denominator) {180 rate <- numerator / denominator181 return(rate)182 }183 ```1841855. **Break long pipes**186 ```r187 # Before (12 operations in one chain)188 result <- data |>189 filter(...) |> mutate(...) |> group_by(...) |>190 summarize(...) |> arrange(...) |> [7 more ops]191192 # After193 clean_data <- data |>194 dplyr::filter(!is.na(value)) |>195 dplyr::mutate(category = categorize(value))196197 summary_stats <- clean_data |>198 dplyr::group_by(category) |>199 dplyr::summarize(mean_val = mean(value))200 ```2012026. **Format and validate**203 ```r204 styler::style_file("analysis.R")205 lintr::lint("analysis.R")206 ```207208**Expected outcome**: Score drops from 60+ to <20209210---211212### Workflow 2: Fix Generic Package Documentation213214**Context**: R package has generic roxygen documentation.215216**Steps**:2172181. **Identify generic patterns**219 ```r220 # Bad221 #' Process Data222 #'223 #' @description This function processes the data.224 #' @param data The data.225 #' @return The result.226 ```2272282. **Make description specific**229 ```r230 # Good231 #' Calculate age-adjusted mortality rates232 #'233 #' Computes mortality rates per 100,000 population, standardized to the234 #' 2000 US Census age distribution using direct standardization.235 ```2362373. **Describe parameter structure**238 ```r239 # Good240 #' @param deaths Data frame with columns `age_group` and `count`.241 #' @param population Data frame with columns `age_group` and `pop_size`.242 ```2432444. **Specify return value**245 ```r246 # Good247 #' @return A tibble with columns:248 #' \describe{249 #' \item{county}{County FIPS code}250 #' \item{rate}{Age-adjusted rate per 100,000}251 #' \item{se}{Standard error of the rate}252 #' }253 ```2542555. **Add realistic examples**256 ```r257 # Good258 #' @examples259 #' counties <- data.frame(260 #' county = c("A", "B"),261 #' deaths = c(150, 200),262 #' population = c(50000, 80000)263 #' )264 #'265 #' adjust_rates(counties, rate_per = 100000)266 #' #> # A tibble: 2 x 3267 #' #> county rate se268 #' #> 1 A 312. 25.4269 #' #> 2 B 258. 18.2270 ```271272**Expected outcome**: Documentation that teaches, not restates273274---275276### Workflow 3: Prepare Package for CRAN277278**Context**: Final checks before CRAN submission.279280**Steps**:2812821. **Run all quality checks**283 ```r284 # Standard checks285 devtools::check()286287 # Anti-slop checks288 lapply(list.files("R", full.names = TRUE), function(f) {289 system(paste("Rscript toolkit/scripts/detect_slop.R", f))290 })291 ```2922932. **Fix documentation**294 - Check all `@param` descriptions are specific295 - Verify `@examples` run and are realistic296 - Ensure `@return` describes structure2972983. **Validate code quality**299 ```r300 # Format all files301 styler::style_dir("R/")302303 # Check lints304 lintr::lint_package()305 ```3063074. **Check CRAN-specific requirements**308 - Use external/posit-skills/r-lib/cran-extrachecks skill309 - Validate URLs in DESCRIPTION and documentation310 - Check examples run in < 5 seconds311312**Expected outcome**: Clean `R CMD check` with no slop patterns313314## Mandatory Rules Summary315316### 1. Namespace Qualification317**ALWAYS use `::` for external packages**318319Exceptions (don't need `::`):320- Base R: `mean()`, `sum()`, `log()`, etc.321- stats: `lm()`, `glm()`, `t.test()`, etc.322- utils: `head()`, `tail()`, `str()`, etc.323324### 2. Explicit Returns325**ALWAYS use `return()` - never implicit**326327### 3. Naming: snake_case328**All objects use `snake_case`**329- Variables: `customer_data` not `customerData` or `df`330- Functions: `calculate_rate` not `calculateRate`331- Arguments: `input_data` not `inputData`332333### 4. Native Pipe334**Prefer `|>` over `%>%`** (unless R < 4.1)335336### 5. No Generic Names337**Never use**: `df`, `data`, `result`, `temp`, `x`, `n` (except standard math notation)338339## Tidyverse Philosophy340341Follow [Tidyverse Style Guide](https://style.tidyverse.org/) as primary reference:3423431. **Design for humans** - Code should be readable and intuitive3442. **Reuse existing data structures** - Work with tibbles and data frames3453. **Compose simple functions with pipes** - Build complexity through composition3464. **Embrace functional programming** - Functions are first-class objects347348See **reference/tidyverse.md** for complete tidyverse conventions.349350## Resources & Advanced Topics351352### Reference Files353354- **[reference/naming.md](reference/naming.md)** - Complete naming conventions and forbidden patterns355- **[reference/tidyverse.md](reference/tidyverse.md)** - Pipe conventions, formatting, ggplot2 standards356- **[reference/documentation.md](reference/documentation.md)** - Roxygen2, vignettes, README quality357- **[reference/statistical-rigor.md](reference/statistical-rigor.md)** - Validation, uncertainty, reproducibility358- **[reference/forbidden-patterns.md](reference/forbidden-patterns.md)** - Complete antipattern catalog359360### Related Skills361362- **external/posit-skills/r-lib/cli** - Error message formatting with cli package363- **external/posit-skills/r-lib/testing** - Test structure and best practices364- **external/posit-skills/r-lib/cran-extrachecks** - CRAN submission requirements365- **text/anti-slop** - For cleaning prose in documentation366367### Tools368369- `styler::style_file()` - Auto-format code370- `lintr::lint()` - Check code quality371- `Rscript toolkit/scripts/detect_slop.R` - Detect AI patterns372373## Integration with Posit Skills374375This skill focuses on **code quality and avoiding generic patterns**.376377Use together with Posit skills for complete coverage:378379| Task | Use This Skill | + Posit Skill |380|------|----------------|---------------|381| Write error messages | r/anti-slop (quality) | + r-lib/cli (structure) |382| Write tests | r/anti-slop (code quality) | + r-lib/testing (test patterns) |383| Prepare for CRAN | r/anti-slop (no slop) | + r-lib/cran-extrachecks (requirements) |384| Document lifecycle | r/anti-slop (doc quality) | + r-lib/lifecycle (deprecation) |