Results for “reward-functions”
50 skillsMore results
reward-function-v410
v4.1.0 reward function redesign to fix overtrading and DSR dominance
3
reward-function-v330
Use when tuning RL reward function weights, fixing reward hacking, or addressing P&L gradient issues
3
reward-function-hold-bias
Fix HOLD bias in RL reward function. Trigger when: (1) model learns to always HOLD, (2) trade rate is too low (<10%), (3) slippage penalty exceeds typical price moves.
3
209-sql-23f1987a
Provides SQL window function examples for ranking, aggregation, lag/lead, value extraction, frame specifications, and advanced analytics.
7 · bundle
cx-incentive-design
Use to design support incentives that improve behaviour without destroying the metric — pairing pay with guardrails, naming gaming modes, and choosing measures that survive Goodhart pressure. Trigger for "incentive plan", "agent bonus scheme", "SPIFF design", "pay for QA score", "what metric should we bonus", CSAT incentives, or reviewing whether a comp change is driving gaming.
1
account-aware-training
Add account state (P&L, win rate, drawdown) to RL observations + drawdown penalty in rewards. Trigger when: (1) model needs account awareness, (2) training should penalize drawdowns, (3) upgrading obs_dim 5300→5600.
3
boost-prompt
Refines task prompts through iterative questioning about scope, deliverables, and constraints, then copies the final markdown to the clipboard.
36.2k
trl-training
Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning) with support for SFT, DPO, GRPO, KTO, RLOO, and reward model training via CLI commands.
10.8k
drive-motivation
Design motivation systems for products and teams using Autonomy, Mastery, and Purpose (AMP) to drive intrinsic engagement.
1.6k · bundle
pump-token-incentives
Volume-based PUMP token reward system with day-indexed epochs, pro-rata distribution, accumulator tracking, and cross-program claiming on Solana.
9
agent-validation-v420
Agent validation overhaul: reward weight overrides, fitness decline gate, pinned data, staged experiments
3
economy-design
Design in-game economies including currencies, progression systems, loot tables, monetization, and resource flow. Use when building the economic backbone of a game. Also trigger for "game economy", "loot table", "drop rates", "currency system", "progression design", "monetization strategy", or "reward system".
0
scarf-model
Diagnoses social threats and rewards across five domains to improve organizational, team, and product experiences.
2
feature-flag-rollout
Use this skill for feature flags, gradual rollout, kill switches, metrics, experiment gates, rollback. Trigger when the task involves programming work related to Feature Flag Rollout, implementation, audits, debugging, strategy, or validation.
1 · bundle
scorecard-marketing
Build quiz and assessment funnels that generate qualified leads through interactive scorecards, covering concept hooks, question design, dynamic results, and automated follow-up sequences.
1.6k · bundle
job-to-be-done-map
Frame jobs, triggers, desired outcomes, anxieties, and product opportunities.
0
fine-tuning-with-trl
Fine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training.
10.4k · bundle
zen
Variable name improvement, function extraction, magic number constants, dead code removal, and code review. For refactoring and PR review — does not change behavior. Don't use for bug/security (Judge), new tests (Radar), architecture (Atlas), or feature implementation (Builder).
3 · bundle
influence-functions-in-deep-learning-arxiv-2002-08484v3
Influence Functions in Deep Learning
6
hook-model
Nir Eyal engagement loop design — trigger, action, variable reward, investment
1 · bundle
tw-review-fold
Classifies code review findings to choose the smallest correct response, preventing one-patch-per-comment churn.
7
review-animations
Reviews animation and motion code against a high craft bar derived from Emil Kowalski's design engineering philosophy. Default to flagging; approval is earned.
5k · bundle
deal-desk
Scores deal margin and risk, routes discount approval to the right human approver, and redlines terms against commercial policy for per-deal review.
20.4k · bundle
functional
Refactors Swift code toward declarative functional programming with immutability, value types, pure functions, transformations, and closure-based APIs.
7
soma
Guides users through participating in the SOMA decentralized training network, covering data submission, model training, reward claiming, and strategic optimization.
32 · bundle
grpo-rl-training
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training
0 · bundle
fine-tuning-with-trl
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.
1 · bundle
stakr-protocol
Interact with the Stakr protocol to create ERC-4626 vaults, add and modify multi-reward staking programs, and manage agent-owned vaults.
1.2k · bundle
autotelic-goal-setting
当设定个人或工作目标,希望目标本身能带来持续的内在动力和满足感时
11 · bundle
team
Defines the architecture and orchestration primitives for multi-agent teams, covering role allocation, conflict resolution, reward routing, and the operational lifecycle of agent collectives.
32
holonic-alignment
当寻求个人行动或项目在更宏大系统中的意义,或需要将日常工作与长远价值连接时
11 · bundle
pump-claims-readonly
Read-only query methods for PumpFun claims — unclaimed token rewards, creator vault balances, volume accumulators, distributable fees, and current-day token previews across Pump and PumpAMM programs.
9
drive-motivation
Design motivation systems using Autonomy, Mastery, and Purpose (AMP) for products and teams. Use when the user mentions "intrinsic motivation", "gamification isnt working", "team incentives", "autonomy", "mastery", "purpose-driven", "employee engagement", or "reward systems". Also trigger when designing onboarding progression systems, fixing broken gamification, or building team structures that sustain high performance. Covers why carrot-and-stick fails and how to build progress systems. For habit-forming product loops, see hooked-ux. For retention behavior design, see improve-retention.
28 · bundle
mariadb-window-functions
Reference for MariaDB window functions and the OVER clause, covering ranking, value navigation, and ordered-set functions with syntax and common pitfalls.
0
grpo-rl-training
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training
1 · bundle