R Anti-Slop: Stop Writing df <- data
When to Use This
Use this for:
- ✓ Any R code leaving your machine (analysis, packages, scripts)
- ✓ AI-generated code review (catches
df, result, missing ::)
- ✓ CRAN submissions (they'll reject generic code anyway)
- ✓ Team code standards
Skip for:
- Quick console experiments (though habits form fast)
- Legacy code you can't touch
- Bioconductor or other style guides that override this
Quick Example
Before (AI Slop):
# Load the library
library(dplyr)
# Read the data
df <- read.csv("data.csv")
# Filter the data
result <- df %>% filter(x > 0)
After (Anti-Slop):
customer_data <- readr::read_csv("data/customers.csv")
active_customers <- customer_data |>
dplyr::filter(status == "active", revenue > 0)
return(active_customers)
What changed:
- ✓ Descriptive names (
customer_data not df)
- ✓ Namespace qualification (
dplyr::, readr::)
- ✓ Native pipe (
|> not %>%)
- ✓ No obvious comments
- ✓ Explicit return
When to Use What
| If you need to... |
Do this |
Details |
| Name variables |
Use snake_case, no df/data/result |
reference/naming.md |
| Call tidyverse functions |
Always use :: (e.g., dplyr::filter()) |
reference/tidyverse.md |
| Return from function |
Always explicit return() statement |
reference/naming.md |
| Write pipe chains |
Use |>, break at 8+ operations |
reference/tidyverse.md |
| Document functions |
Specific @param, @return, no circular text |
reference/documentation.md |
| Handle missing data |
Explicit strategy + report data loss |
reference/statistical-rigor.md |
| Validate data |
Check assumptions with stopifnot() |
reference/statistical-rigor.md |
| Format code |
Use styler::style_file() |
reference/tidyverse.md |
| Check code quality |
Use lintr::lint() |
reference/tidyverse.md |
Core Workflow
5-Step Quality Check
Namespace qualification - All external functions use ::
# Good
dplyr::filter(data, x > 0)
# Bad
filter(data, x > 0)
Explicit returns - Every function has return()
# Good
my_function <- function(x) {
result <- x + 1
return(result)
}
# Bad
my_function <- function(x) {
x + 1
}
Naming conventions - All objects use snake_case
# Good
customer_lifetime_value <- calculate_clv(data)
# Bad
df <- calculate_clv(data)
customerLifetimeValue <- calculate_clv(data)
Documentation quality - No generic descriptions
# Good
#' @param deaths Data frame with `age_group` and `count` columns
# Bad
#' @param data The data
Code formatting - Run styler and lintr
styler::style_file("script.R")
lintr::lint("script.R")
Quick Reference Checklist
Before committing R code, verify:
Common Workflows
Workflow 1: Clean Up AI-Generated R Script
Context: AI generated an analysis script with generic patterns.
Steps:
Run detection script
Rscript toolkit/scripts/detect_slop.R analysis.R --verbose
Fix high-priority issues first
# Replace df, data, result with descriptive names
# Before
df <- readr::read_csv("data.csv")
result <- df %>% filter(x > 0)
# After
customer_data <- readr::read_csv("data/customers.csv")
active_customers <- customer_data |> dplyr::filter(status == "active")
Add namespace qualification
# Before
data %>% filter(x > 0) %>% summarize(mean(y))
# After
data |>
dplyr::filter(x > 0) |>
dplyr::summarize(mean_y = mean(y))
Add explicit returns
# Before
calculate_rate <- function(numerator, denominator) {
numerator / denominator
}
# After
calculate_rate <- function(numerator, denominator) {
rate <- numerator / denominator
return(rate)
}
Break long pipes
# Before (12 operations in one chain)
result <- data |>
filter(...) |> mutate(...) |> group_by(...) |>
summarize(...) |> arrange(...) |> [7 more ops]
# After
clean_data <- data |>
dplyr::filter(!is.na(value)) |>
dplyr::mutate(category = categorize(value))
summary_stats <- clean_data |>
dplyr::group_by(category) |>
dplyr::summarize(mean_val = mean(value))
Format and validate
styler::style_file("analysis.R")
lintr::lint("analysis.R")
Expected outcome: Score drops from 60+ to <20
Workflow 2: Fix Generic Package Documentation
Context: R package has generic roxygen documentation.
Steps:
Identify generic patterns
# Bad
#' Process Data
#'
#' @description This function processes the data.
#' @param data The data.
#' @return The result.
Make description specific
# Good
#' Calculate age-adjusted mortality rates
#'
#' Computes mortality rates per 100,000 population, standardized to the
#' 2000 US Census age distribution using direct standardization.
Describe parameter structure
# Good
#' @param deaths Data frame with columns `age_group` and `count`.
#' @param population Data frame with columns `age_group` and `pop_size`.
Specify return value
# Good
#' @return A tibble with columns:
#' \describe{
#' \item{county}{County FIPS code}
#' \item{rate}{Age-adjusted rate per 100,000}
#' \item{se}{Standard error of the rate}
#' }
Add realistic examples
# Good
#' @examples
#' counties <- data.frame(
#' county = c("A", "B"),
#' deaths = c(150, 200),
#' population = c(50000, 80000)
#' )
#'
#' adjust_rates(counties, rate_per = 100000)
#' #> # A tibble: 2 x 3
#' #> county rate se
#' #> 1 A 312. 25.4
#' #> 2 B 258. 18.2
Expected outcome: Documentation that teaches, not restates
Workflow 3: Prepare Package for CRAN
Context: Final checks before CRAN submission.
Steps:
Run all quality checks
# Standard checks
devtools::check()
# Anti-slop checks
lapply(list.files("R", full.names = TRUE), function(f) {
system(paste("Rscript toolkit/scripts/detect_slop.R", f))
})
Fix documentation
- Check all
@param descriptions are specific
- Verify
@examples run and are realistic
- Ensure
@return describes structure
Validate code quality
# Format all files
styler::style_dir("R/")
# Check lints
lintr::lint_package()
Check CRAN-specific requirements
- Validate URLs in DESCRIPTION and documentation
- Check examples run in < 5 seconds
- Verify package structure meets CRAN standards
Expected outcome: Clean R CMD check with no slop patterns
Mandatory Rules Summary
1. Namespace Qualification
ALWAYS use :: for external packages
Exceptions (don't need ::):
- Base R:
mean(), sum(), log(), etc.
- stats:
lm(), glm(), t.test(), etc.
- utils:
head(), tail(), str(), etc.
2. Explicit Returns
ALWAYS use return() - never implicit
3. Naming: snake_case
All objects use snake_case
- Variables:
customer_data not customerData or df
- Functions:
calculate_rate not calculateRate
- Arguments:
input_data not inputData
4. Native Pipe
Prefer |> over %>% (unless R < 4.1)
5. No Generic Names
Never use: df, data, result, temp, x, n (except standard math notation)
Tidyverse Philosophy
Follow Tidyverse Style Guide as primary reference:
- Design for humans - Code should be readable and intuitive
- Reuse existing data structures - Work with tibbles and data frames
- Compose simple functions with pipes - Build complexity through composition
- Embrace functional programming - Functions are first-class objects
See reference/tidyverse.md for complete tidyverse conventions.
Resources & Advanced Topics
Reference Files
- reference/naming.md - Complete naming conventions and forbidden patterns
- reference/tidyverse.md - Pipe conventions, formatting, ggplot2 standards
- reference/documentation.md - Roxygen2, vignettes, README quality
- reference/statistical-rigor.md - Validation, uncertainty, reproducibility
- reference/forbidden-patterns.md - Complete antipattern catalog
Related Skills
- text/anti-slop - For cleaning prose in documentation
- quarto/anti-slop - For cleaning vignettes and documentation
Tools
styler::style_file() - Auto-format code
lintr::lint() - Check code quality
Rscript toolkit/scripts/detect_slop.R - Detect AI patterns
Integration with Posit Skills
This skill focuses on code quality and avoiding generic patterns.
Use together with Posit skills for complete coverage:
| Task |
Use This Skill |
+ Posit Skill |
| Write error messages |
r/anti-slop (quality) |
+ r-lib/cli (structure) |
| Write tests |
r/anti-slop (code quality) |
+ r-lib/testing (test patterns) |
| Prepare for CRAN |
r/anti-slop (no slop) |
+ r-lib/cran-extrachecks (requirements) |
| Document lifecycle |
r/anti-slop (doc quality) |
+ r-lib/lifecycle (deprecation) |
1---2name: r-anti-slop3description: Enforce production-quality R code standards. Prevents generic AI patterns through namespace qualification, explicit returns, and tidyverse conventions. Use when writing or reviewing R code for data analysis or packages.4---5
6# R Anti-Slop: Stop Writing `df <- data`
7
8## When to Use This
9
10Use this for:
11- ✓ Any R code leaving your machine (analysis, packages, scripts)
12- ✓ AI-generated code review (catches `df`, `result`, missing `::`)
13- ✓ CRAN submissions (they'll reject generic code anyway)
14- ✓ Team code standards
15
16Skip for:
17- Quick console experiments (though habits form fast)
18- Legacy code you can't touch
19- Bioconductor or other style guides that override this
20
21## Quick Example
22
23**Before (AI Slop)**:
24```r
25# Load the library
26library(dplyr)
27
28# Read the data
29df <- read.csv("data.csv")
30
31# Filter the data
32result <- df %>% filter(x > 0)
33```
34
35**After (Anti-Slop)**:
36```r
37customer_data <- readr::read_csv("data/customers.csv")
38
39active_customers <- customer_data |>
40 dplyr::filter(status == "active", revenue > 0)
41
42return(active_customers)
43```
44
45**What changed**:
46- ✓ Descriptive names (`customer_data` not `df`)
47- ✓ Namespace qualification (`dplyr::`, `readr::`)
48- ✓ Native pipe (`|>` not `%>%`)
49- ✓ No obvious comments
50- ✓ Explicit return
51
52## When to Use What
53
54| If you need to... | Do this | Details |
55|-------------------|---------|---------|
56| Name variables | Use `snake_case`, no `df`/`data`/`result` | reference/naming.md |
57| Call tidyverse functions | Always use `::` (e.g., `dplyr::filter()`) | reference/tidyverse.md |
58| Return from function | Always explicit `return()` statement | reference/naming.md |
59| Write pipe chains | Use `\|>`, break at 8+ operations | reference/tidyverse.md |
60| Document functions | Specific `@param`, `@return`, no circular text | reference/documentation.md |
61| Handle missing data | Explicit strategy + report data loss | reference/statistical-rigor.md |
62| Validate data | Check assumptions with `stopifnot()` | reference/statistical-rigor.md |
63| Format code | Use `styler::style_file()` | reference/tidyverse.md |
64| Check code quality | Use `lintr::lint()` | reference/tidyverse.md |
65
66## Core Workflow
67
68### 5-Step Quality Check
69
701. **Namespace qualification** - All external functions use `::`
71 ```r
72 # Good
73 dplyr::filter(data, x > 0)
74 # Bad
75 filter(data, x > 0)
76 ```
77
782. **Explicit returns** - Every function has `return()`
79 ```r
80 # Good
81 my_function <- function(x) {
82 result <- x + 1
83 return(result)
84 }
85 # Bad
86 my_function <- function(x) {
87 x + 1
88 }
89 ```
90
913. **Naming conventions** - All objects use `snake_case`
92 ```r
93 # Good
94 customer_lifetime_value <- calculate_clv(data)
95 # Bad
96 df <- calculate_clv(data)
97 customerLifetimeValue <- calculate_clv(data)
98 ```
99
1004. **Documentation quality** - No generic descriptions
101 ```r
102 # Good
103 #' @param deaths Data frame with `age_group` and `count` columns
104 # Bad
105 #' @param data The data
106 ```
107
1085. **Code formatting** - Run styler and lintr
109 ```r
110 styler::style_file("script.R")
111 lintr::lint("script.R")
112 ```
113
114## Quick Reference Checklist
115
116Before committing R code, verify:
117
118- [ ] All external functions qualified with `::`
119- [ ] All functions have explicit `return()`
120- [ ] All objects use `snake_case`
121- [ ] No generic names (`df`, `data`, `result`, `temp`)
122- [ ] Pipes (`|>`) have space before, end lines
123- [ ] Long pipelines (>8 ops) broken into named steps
124- [ ] Complex operations have WHY comments
125- [ ] Data validated after transformations
126- [ ] Seeds set before random operations
127- [ ] Uncertainty reported (SE, CI) for statistical models
128- [ ] No `attach()` calls
129- [ ] No right-hand assignment (`->`)
130- [ ] Roxygen documentation is specific
131- [ ] Examples are realistic and run
132
133## Common Workflows
134
135### Workflow 1: Clean Up AI-Generated R Script
136
137**Context**: AI generated an analysis script with generic patterns.
138
139**Steps**:
140
1411. **Run detection script**
142 ```bash
143 Rscript toolkit/scripts/detect_slop.R analysis.R --verbose
144 ```
145
1462. **Fix high-priority issues first**
147 ```r
148 # Replace df, data, result with descriptive names
149 # Before
150 df <- readr::read_csv("data.csv")
151 result <- df %>% filter(x > 0)
152
153 # After
154 customer_data <- readr::read_csv("data/customers.csv")
155 active_customers <- customer_data |> dplyr::filter(status == "active")
156 ```
157
1583. **Add namespace qualification**
159 ```r
160 # Before
161 data %>% filter(x > 0) %>% summarize(mean(y))
162
163 # After
164 data |>
165 dplyr::filter(x > 0) |>
166 dplyr::summarize(mean_y = mean(y))
167 ```
168
1694. **Add explicit returns**
170 ```r
171 # Before
172 calculate_rate <- function(numerator, denominator) {
173 numerator / denominator
174 }
175
176 # After
177 calculate_rate <- function(numerator, denominator) {
178 rate <- numerator / denominator
179 return(rate)
180 }
181 ```
182
1835. **Break long pipes**
184 ```r
185 # Before (12 operations in one chain)
186 result <- data |>
187 filter(...) |> mutate(...) |> group_by(...) |>
188 summarize(...) |> arrange(...) |> [7 more ops]
189
190 # After
191 clean_data <- data |>
192 dplyr::filter(!is.na(value)) |>
193 dplyr::mutate(category = categorize(value))
194
195 summary_stats <- clean_data |>
196 dplyr::group_by(category) |>
197 dplyr::summarize(mean_val = mean(value))
198 ```
199
2006. **Format and validate**
201 ```r
202 styler::style_file("analysis.R")
203 lintr::lint("analysis.R")
204 ```
205
206**Expected outcome**: Score drops from 60+ to <20
207
208---
209
210### Workflow 2: Fix Generic Package Documentation
211
212**Context**: R package has generic roxygen documentation.
213
214**Steps**:
215
2161. **Identify generic patterns**
217 ```r
218 # Bad
219 #' Process Data
220 #'
221 #' @description This function processes the data.
222 #' @param data The data.
223 #' @return The result.
224 ```
225
2262. **Make description specific**
227 ```r
228 # Good
229 #' Calculate age-adjusted mortality rates
230 #'
231 #' Computes mortality rates per 100,000 population, standardized to the
232 #' 2000 US Census age distribution using direct standardization.
233 ```
234
2353. **Describe parameter structure**
236 ```r
237 # Good
238 #' @param deaths Data frame with columns `age_group` and `count`.
239 #' @param population Data frame with columns `age_group` and `pop_size`.
240 ```
241
2424. **Specify return value**
243 ```r
244 # Good
245 #' @return A tibble with columns:
246 #' \describe{
247 #' \item{county}{County FIPS code}
248 #' \item{rate}{Age-adjusted rate per 100,000}
249 #' \item{se}{Standard error of the rate}
250 #' }
251 ```
252
2535. **Add realistic examples**
254 ```r
255 # Good
256 #' @examples
257 #' counties <- data.frame(
258 #' county = c("A", "B"),
259 #' deaths = c(150, 200),
260 #' population = c(50000, 80000)
261 #' )
262 #'
263 #' adjust_rates(counties, rate_per = 100000)
264 #' #> # A tibble: 2 x 3
265 #' #> county rate se
266 #' #> 1 A 312. 25.4
267 #' #> 2 B 258. 18.2
268 ```
269
270**Expected outcome**: Documentation that teaches, not restates
271
272---
273
274### Workflow 3: Prepare Package for CRAN
275
276**Context**: Final checks before CRAN submission.
277
278**Steps**:
279
2801. **Run all quality checks**
281 ```r
282 # Standard checks
283 devtools::check()
284
285 # Anti-slop checks
286 lapply(list.files("R", full.names = TRUE), function(f) {
287 system(paste("Rscript toolkit/scripts/detect_slop.R", f))
288 })
289 ```
290
2912. **Fix documentation**
292 - Check all `@param` descriptions are specific
293 - Verify `@examples` run and are realistic
294 - Ensure `@return` describes structure
295
2963. **Validate code quality**
297 ```r
298 # Format all files
299 styler::style_dir("R/")
300
301 # Check lints
302 lintr::lint_package()
303 ```
304
3054. **Check CRAN-specific requirements**
306 - Validate URLs in DESCRIPTION and documentation
307 - Check examples run in < 5 seconds
308 - Verify package structure meets CRAN standards
309
310**Expected outcome**: Clean `R CMD check` with no slop patterns
311
312## Mandatory Rules Summary
313
314### 1. Namespace Qualification
315**ALWAYS use `::` for external packages**
316
317Exceptions (don't need `::`):
318- Base R: `mean()`, `sum()`, `log()`, etc.
319- stats: `lm()`, `glm()`, `t.test()`, etc.
320- utils: `head()`, `tail()`, `str()`, etc.
321
322### 2. Explicit Returns
323**ALWAYS use `return()` - never implicit**
324
325### 3. Naming: snake_case
326**All objects use `snake_case`**
327- Variables: `customer_data` not `customerData` or `df`
328- Functions: `calculate_rate` not `calculateRate`
329- Arguments: `input_data` not `inputData`
330
331### 4. Native Pipe
332**Prefer `|>` over `%>%`** (unless R < 4.1)
333
334### 5. No Generic Names
335**Never use**: `df`, `data`, `result`, `temp`, `x`, `n` (except standard math notation)
336
337## Tidyverse Philosophy
338
339Follow [Tidyverse Style Guide](https://style.tidyverse.org/) as primary reference:
340
3411. **Design for humans** - Code should be readable and intuitive
3422. **Reuse existing data structures** - Work with tibbles and data frames
3433. **Compose simple functions with pipes** - Build complexity through composition
3444. **Embrace functional programming** - Functions are first-class objects
345
346See **reference/tidyverse.md** for complete tidyverse conventions.
347
348## Resources & Advanced Topics
349
350### Reference Files
351
352- **[reference/naming.md](reference/naming.md)** - Complete naming conventions and forbidden patterns
353- **[reference/tidyverse.md](reference/tidyverse.md)** - Pipe conventions, formatting, ggplot2 standards
354- **[reference/documentation.md](reference/documentation.md)** - Roxygen2, vignettes, README quality
355- **[reference/statistical-rigor.md](reference/statistical-rigor.md)** - Validation, uncertainty, reproducibility
356- **[reference/forbidden-patterns.md](reference/forbidden-patterns.md)** - Complete antipattern catalog
357
358### Related Skills
359
360- **text/anti-slop** - For cleaning prose in documentation
361- **quarto/anti-slop** - For cleaning vignettes and documentation
362
363### Tools
364
365- `styler::style_file()` - Auto-format code
366- `lintr::lint()` - Check code quality
367- `Rscript toolkit/scripts/detect_slop.R` - Detect AI patterns
368
369## Integration with Posit Skills
370
371This skill focuses on **code quality and avoiding generic patterns**.
372
373Use together with Posit skills for complete coverage:
374
375| Task | Use This Skill | + Posit Skill |
376|------|----------------|---------------|
377| Write error messages | r/anti-slop (quality) | + r-lib/cli (structure) |
378| Write tests | r/anti-slop (code quality) | + r-lib/testing (test patterns) |
379| Prepare for CRAN | r/anti-slop (no slop) | + r-lib/cran-extrachecks (requirements) |
380| Document lifecycle | r/anti-slop (doc quality) | + r-lib/lifecycle (deprecation) |