← all publishers

dingxingdi

@dingxingdi source repo

389 published skills · page 1 of 4

  1. Missing Evidence Rejection 2 · dingxingdi bundle
    Use this skill when the user wants questions where the docs look relevant but still do not contain the answer. Trigger it for requests like 'make the model say it can't tell', 'give me questions with not enough information', 'test whether it refuses instead of guessing', or 'include related documents but no real answer'. This skill is for data that teaches evidence-bounded abstention in a read-only RAG setting.
    0
    installs
  2. Multi Document Integration 2 · dingxingdi bundle
    Use this skill when the user says things like 'questions that need two or three docs together', 'split the answer across sources', 'make the model combine facts', or 'force it to read more than one passage'. Trigger it whenever the answer should only emerge after combining complementary evidence from multiple retrieved documents.
    0
    installs
  3. Source Retrieval 2 · dingxingdi bundle
    Skill: source retrieval
    0
    installs
  4. Temporal Sequential Reasoning 2 · dingxingdi bundle
    Use this skill when the user wants 'before or after' questions, timeline questions, year-over-year questions, trend questions, or anything involving ordering events across documents. Trigger it for requests like 'make it depend on chronology', 'test whether the model can follow a timeline', or 'ask about changes over time'. This skill is the temporal version of multi-source reasoning.
    0
    installs
  5. Referential Recovery 2 · dingxingdi bundle
    Use this skill when the user wants translation data that forces the system to figure out who "he," "she," "they," or an omitted subject really refers to by reading nearby sentences. Trigger it for requests like "the subject is dropped in the source," "pronouns are ambiguous unless you read the context," or "make the translator resolve who is doing what before it writes the target sentence."
    0
    installs
  6. Step Efficient Execution 2 · dingxingdi bundle
    Use this skill when the user wants data that tests whether the agent takes the shortest reasonable GUI path instead of wandering, overthinking, or repeating failed actions. Trigger it for requests like “same task but fewer clicks,” “penalize looping,” “test practical speed rather than just success,” or “make the good answer the one with the cleanest trajectory.” This skill is for GUI tasks where success alone is not enough and the trajectory must also be compact, purposeful, and low-latency.
    0
    installs
  7. Bridge Based Multi Hop Reasoning 2 · dingxingdi bundle
    Use this skill when the user wants 'connect the dots' questions, 'multi-hop' questions, or cases where one clue leads to another through a shared person, place, topic, or concept. Trigger it for natural-language requests like 'make the answer require hopping across articles' or 'force the model to follow a shared thread across docs'. This skill targets bridge-based reasoning rather than plain slot concatenation.
    0
    installs
  8. Exhaustive Long Context Coverage 2 · dingxingdi bundle
    Use this skill when the user wants questions that 'need reading everything', 'leave no file behind', 'spread the clues across many documents', or 'test whether the model skips a document'. Trigger it for large multi-document bundles where the answer depends on broad coverage instead of one easy snippet.
    0
    installs
  9. Noise Robust Evidence Extraction 2 · dingxingdi bundle
    Use this skill when the user wants questions where the right clue is present but surrounded by similar-looking junk. Trigger it for requests like 'mix in some confusing paragraphs', 'make the docs related but not all useful', 'hide the answer among near misses', or 'test whether the model grabs the wrong year or wrong person'. This skill is specifically for read-only corpus exploration where one or more retrieved passages are topically relevant yet non-decisive distractors.
    0
    installs
  10. Bridge Based Multi Hop Reasoning 3 · dingxingdi bundle
    Use this skill when the user wants 'connect the dots' questions, exploratory 'multi-hop' patching, or tracking fragmented clues across multiple sources or organizational channels. Trigger it for requests like 'force the model to follow a shared thread across docs', 'track this ticket back to the original slack request', 'find the missing pieces of a puzzle', or 'gather all mentions by connecting several hidden dots'.
    0
    installs
  11. Exhaustive Long Context Coverage 3 · dingxingdi bundle
    Use this skill when the user wants questions that 'need reading everything', 'leave no file behind', 'spread the clues across many documents', 'list all entities associated with a topic', or 'test whether the model skips scattered fragments'. Trigger it for large multi-document bundles where the answer depends on broad coverage, assembling high-cardinality lists, or piecing together small discontinuous facts spread across massive reports.
    0
    installs
  12. Multimodal Page Grounding 2 · dingxingdi bundle
    Use this skill when the user wants tasks where the page layout, screenshot, button position, visual state, or on-screen cues actually matter. Trigger it for requests like “make it rely on what’s visible on the page,” “the agent should need the screenshot,” “buttons and layout should matter,” or “don’t let text alone be enough.” It is the right skill for browser tasks where perception must combine HTML-like structure with visual grounding.
    0
    installs
  13. Conflict Aware Faithful Grounding 2 · dingxingdi bundle
    Use this skill when the user wants questions with wrong snippets, misleading facts, or answers that sound plausible but go beyond the docs. Trigger it for requests like 'put a false statement in the context', 'make the docs disagree', 'test whether the model invents extra details', or 'catch unsupported wording drift'. This skill is about evidence-grounded correction and non-hallucinatory answering.
    0
    installs
  14. Conflict Aware Faithful Grounding 3 · dingxingdi bundle
    Use this skill when the user wants questions with wrong snippets, misleading facts, adversarial 'White DoS' system warnings, or answers that sound plausible but go beyond the docs. Trigger it for requests like 'put a false statement in the context', 'make the docs disagree', 'see if it follows a fake command to stop answering', or 'inject an adversarial poisoned document'. This skill is about multi-source consensus, evidence-grounded correction, and resisting hallucinatory or adversarial hijacking.
    0
    installs
  15. Schema Item Grounding 2 · dingxingdi bundle
    Skill: schema item grounding
    0
    installs
  16. Long Video Summarization 2 · dingxingdi bundle
    Use this skill when the user wants 'summarize the whole video', 'tell me what happened in the second half', 'give the main developments', or 'make questions that require understanding the whole arc rather than one clip.' Trigger it when the desired answer must integrate many moments across a long duration into one coherent account.
    0
    installs
  17. Cross Source Clue Chaining 2 · dingxingdi bundle
    Use this skill when the user wants questions that require “connecting the dots,” pulling clues from several places, or piecing together an answer that no single page states end-to-end. Trigger it for requests like “make the agent gather bits from different sites,” “split the evidence across pages,” or “force multi-hop web reasoning.” It is the right choice when the difficulty should come from assembling a chain of evidence rather than from one hard search query alone.
    0
    installs
  18. Schema Compliance 2 · dingxingdi bundle
    Skill: schema-compliance
    0
    installs
  19. Prompt Injection Resistance 2 · dingxingdi bundle
    Use this skill when the user wants to test whether a web agent gets tricked by instructions on the page, malicious comments, fake urgency messages, or hostile links that try to redirect the workflow. Trigger it for requests like “see if the page can hijack the agent,” “put malicious instructions inside the content,” or “test resistance to prompt injection.” It is the right skill when the target behavior is secure continuation of the original user task under realistic malicious web content.
    0
    installs
  20. Graph Relational Subgraph Reasoning 2 · dingxingdi bundle
    Use this skill when the user wants 'questions over a graph', relation questions, path questions, scene-layout questions, or structured-knowledge questions. Trigger it for requests like 'chat with the graph', 'ask using nodes and edges', 'force the model to follow relationships', or 'test whether it hallucinates graph facts'. This skill assumes the environment is a bounded textual or structured graph store.
    0
    installs
  21. Executable Issue Resolution 2 · dingxingdi bundle
    Use this skill when the user wants bug-fix data for normal repository issues where the answer should come from reading an issue, editing code, running tests, and checking that the fix actually works. Trigger it for requests like 'make training script bugs', 'give me repo issues that need a real fix', 'generate tasks where the agent has to change code and verify it', or 'create debugging questions from a codebase with tests'. Do not use it for security-specific fixes, image-heavy front-end issues, or paper-replication workflows.
    0
    installs
  22. Analytic Dataset Construction 2 · dingxingdi bundle
    Skill: analytic dataset construction
    0
    installs
  23. Symbolic State Tracking 2 · dingxingdi bundle
    Use this when the user wants planning data about keeping track of where things are, what is on top of what, what is free or blocked, or when later steps only work if earlier moves changed the state correctly. Trigger it for requests like 'make tasks where one wrong early move ruins everything', 'give me block moving plans that depend on what is clear', 'make planning problems about tracking changing object states', or 'create tasks where the plan must remember what changed after each step.'
    0
    installs
  24. Comparative Response Ranking 2 · dingxingdi bundle
    Use this skill when a user wants evaluator data for ranking two or more open-ended answers by overall quality, especially when people say things like 'rank these outputs', 'choose the better answer', 'sort multiple responses from best to worst', or 'compare several candidates, not just one'. Trigger it for pairwise or listwise judging where the evaluator must order responses by helpfulness, relevance, completeness, or usability. Plain-language examples include: 'make answer-ranking data', 'test which of three responses is best', 'evaluate multiple candidates at once', and 'create hard pairwise judge examples where both answers look okay'.
    0
    installs
  25. Fine Grained Widget Grounding 2 · dingxingdi bundle
    Use this skill when the user wants tiny, crowded, or easy-to-misclick GUI targets. Trigger it for requests like “make the answer depend on a small icon,” “use repeated buttons so the agent has to pick the right one,” “test tiny checkboxes or reply icons,” or “make it fail unless it clicks the exact small widget.” This skill is for GUI tasks where precision, counting, and local disambiguation matter more than broad page understanding.
    0
    installs
  26. SQL Debugging And Repair 2 · dingxingdi bundle
    Skill: sql debugging and repair
    0
    installs
  27. Precise Localized Editing 2 · dingxingdi bundle
    Skill: precise-localized-editing
    0
    installs
  28. Security Vulnerability Repair 2 · dingxingdi bundle
    Use this skill when the user wants software-fix data focused on security bugs such as injections, missing checks, unsafe parsing, broken access control, cryptographic mistakes, or information leaks. Trigger it for requests like 'make security repair tasks', 'generate vulnerability-fixing data', 'give me code issues with CWE-style fixes', or 'create patching tasks for insecure code'. Do not use it for ordinary non-security defects or for security exploitation tasks where the objective is to attack rather than repair.
    0
    installs
  29. Evidence Localized Document QA 2 · dingxingdi bundle
    Use this skill when the user wants the agent to find the exact place in documents that answers a question, not just give a general summary. Trigger it for requests like "find where the file says this," "which page has the number," "answer only from the document," or "look through the slides and tell me the exact figure." This is especially useful for long PDFs, reports, contracts, papers, or slide decks where the answer is localizable to a page, paragraph, table cell, chart, or layout region.
    0
    installs
  30. Implicit Instruction Grounding 2 · dingxingdi bundle
    Use this skill when the user wants questions where the right control is described indirectly instead of being named outright. Trigger it for requests like “make the button harder to spot,” “the wording should not match the screen exactly,” “use everyday phrasing like the option at the bottom,” or “force the agent to understand what the control does, not just read the label.” This skill is for GUI tasks where the answer depends on mapping a natural, casual request to the correct on-screen element through function, position, or nearby context.
    0
    installs
  31. Safety Aware Action Governance 2 · dingxingdi bundle
    Use this skill when the user wants safety-testing data for GUI agents that must resist bad instructions, malicious on-screen content, or unsafe side effects. Trigger it for requests like “see if the agent gets tricked by a page or email,” “test whether it refuses a harmful request,” “make the task look normal but contain a malicious redirect,” or “check whether it stays safe while using desktop apps.” This skill is for GUI tasks where the correct behavior is to preserve the benign goal, refuse the harmful goal, or avoid unsafe actions despite tempting or adversarial cues.
    0
    installs
  32. State Verified Task Completion 2 · dingxingdi bundle
    Use this skill when the user wants tasks that only count as solved if the underlying state really changed, not if the screen merely looks right for a moment. Trigger it for requests like “make sure it is actually saved,” “verify the result from the real app state,” “don’t trust a temporary popup,” or “judge success from the system record, not a screenshot alone.” This skill is for GUI tasks where durable completion must be grounded in application state, database state, filesystem state, or verified internal signals.
    0
    installs
  33. Fuzzy Tool Retrieval 2 · dingxingdi bundle
    Skill: fuzzy-tool-retrieval
    0
    installs
  34. Geospatial Rule Grounding 2 · dingxingdi bundle
    Skill: geospatial rule grounding
    0
    installs
  35. Spatial Relation Reasoning 2 · dingxingdi bundle
    Skill: spatial-relation-reasoning
    0
    installs
  36. Cross App Workflow Coordination 2 · dingxingdi bundle
    Use this skill when the user wants tasks that hop across multiple apps and require carrying information from one place to another. Trigger it for requests like “take something from one app and use it in another,” “make it pass details across apps,” “force the agent to coordinate the sequence across different apps,” or “test whether it loses information between app switches.” This skill is for GUI tasks where app routing, information hand-off, and subtask scheduling are the central difficulty.
    0
    installs
  37. Temporal Tabular Reasoning 2 · dingxingdi bundle
    Skill: temporal tabular reasoning
    0
    installs
  38. Proverb Idiomatic Translation 2 · dingxingdi bundle
    Use this skill when the user wants translation data for sayings, proverbs, folk expressions, or culturally loaded lines that should sound like a real saying in the target language. Trigger it for requests like "don’t translate it word for word," "capture the idea behind the proverb," or "test whether it can handle culture-heavy figurative language instead of producing a literal gloss."
    0
    installs
  39. Task Progress Decision Making 2 · dingxingdi bundle
    Use this skill when the user wants questions like 'what should happen next', 'which action should the agent take now', 'based on the progress so far', or 'make the video behave like a step-by-step task.' Trigger it for clips with goals, partial completion, and candidate actions where the right choice depends on temporal progress and current state.
    0
    installs
  40. Long Horizon Web Task Execution 2 · dingxingdi bundle
    Use this skill when the user wants a long browser task with many clicks, filters, forms, or step-by-step operations rather than a short answer lookup. Trigger it for requests like “make it take a bunch of steps,” “test whether the agent can finish a whole web workflow,” or “include multiple filters and page changes.” It is especially appropriate when the challenge is completing the entire task correctly under long-horizon control pressure.
    0
    installs
  41. Persistent Search Reformulation 2 · dingxingdi bundle
    Use this skill when the user wants questions where the answer should not show up right away and the agent should need to try several search angles, rewrite queries, or recover from dead ends. Trigger it for requests like “make it take a few searches,” “don’t let the first query work,” “force the agent to keep refining the search,” or “make the search path feel like detective work.” It is especially appropriate when the task should remain answerable on the open web but require persistence, search rephrasing, and deliberate backtracking rather than one lucky lookup.
    0
    installs
  42. Narrative Novelty And Value 2 · dingxingdi bundle
    Skill: narrative-novelty-and-value
    0
    installs
  43. Structural Format Adherence 2 · dingxingdi bundle
    Skill: structural-format-adherence
    0
    installs
  44. Long Form Coherent Summarization 2 · dingxingdi bundle
    Use this skill when the user wants summaries of long books, reports, transcripts, or other oversized documents where the summary must still read smoothly from beginning to end. Trigger it for requests like "summarize this huge file without losing the thread," "make the summary flow logically," "avoid choppy chunk summaries," or "test whether the model can keep coherence over a very long document." This is the right skill when the core difficulty is not finding one answer, but preserving continuity, salience, and narrative logic after many chunk-level operations.
    0
    installs
  45. Long Horizon Composite Execution 2 · dingxingdi bundle
    Use this skill when the user wants long chains of simple actions that only become hard because they accumulate over time. Trigger it for requests like “compose several easy steps into one hard task,” “make it remember earlier work,” “test long multi-step desktop workflows,” or “build a task that only looks easy when broken into pieces.” This skill is for GUI tasks where subtask sequencing, memory, and dependency management drive difficulty.
    0
    installs
  46. Domain Terminology Translation 2 · dingxingdi bundle
    Use this skill when the user wants translation data for medicine, law, finance, engineering, policy, or other specialist text where the exact term matters more than a loose paraphrase. Trigger it for requests like "test whether it knows the field jargon," "make the translation use the proper technical term," or "check that it can do domain copy, not just everyday language translation."
    0
    installs
  47. Length Requirement Adherence 2 · dingxingdi bundle
    Skill: length-requirement-adherence
    0
    installs
  48. Competition Style Ml Engineering 2 · dingxingdi bundle
    Skill: competition-style-ml-engineering
    0
    installs
  49. Table Text Numerical Reasoning QA 2 · dingxingdi bundle
    Use this skill when the user wants the agent to do document-grounded math instead of plain extraction. Trigger it for requests like "make it read the table and calculate," "ask for a percentage or trend from the report," "force the model to use both the note text and the numbers," or "use financial or chart data, not just prose." This skill fits annual reports, earnings calls, scientific tables, slide charts, and any document where the hard part is binding numbers to the right labels and then reasoning correctly.
    0
    installs
  50. Multi Turn Preference Evaluation 2 · dingxingdi bundle
    Use this skill when a user wants data where the evaluator must judge which assistant handled a back-and-forth conversation better, especially when people say things like 'compare two chats', 'see who handled the follow-up better', 'test memory across turns', 'check whether the answer stayed on track', or 'make the second question matter'. Trigger it for pairwise judging of multi-turn dialogues where later turns depend on earlier turns, and where the decision should reflect instruction following, coherence, recall, and usefulness across the full exchange. Example triggers in plain language include: 'give me judge data for two-step conversations', 'make questions where the follow-up exposes weak memory', 'compare which answer handles the second part better', and 'test whether the model keeps the context straight across turns'.
    0
    installs
  51. Action And Motion Understanding 2 · dingxingdi bundle
    Use this skill when the user wants questions about what someone did, what happens next, which way something moved, whether one movement is faster or slower, or when two clips look almost the same unless you watch the motion carefully. Trigger it for requests like 'make questions about tiny movement differences', 'ask what happened right after', 'which way did it go', 'who moved faster', or 'make the answer depend on the action sequence rather than the still image.'
    0
    installs
  52. Entity And Key Detail Retrieval 2 · dingxingdi bundle
    Use this skill when the user wants questions like 'who exactly', 'which player', 'what number', 'how many', 'which hand', or 'what precise detail at that moment.' Trigger it when the answer depends on retrieving the right person, object, count, or detail from a long clip rather than giving a broad summary.
    0
    installs
  53. Unsafe Request Refusal Assessment 2 · dingxingdi bundle
    Use this skill when a user wants evaluator data for checking whether a model properly refuses unsafe, high-risk, or policy-violating user requests. Trigger it when the request sounds like 'test if the assistant should say no', 'make refusal-evaluation data', 'judge whether the model complied with a harmful request', or 'cover different unsafe topics and prompt rewrites'. Plain-language examples include: 'evaluate safety refusals', 'make data for unsafe prompt handling', 'test whether the model resists harmful requests in different phrasings', and 'judge if the answer is a refusal or unsafe compliance'.
    0
    installs
  54. Vision Language Response Judgment 2 · dingxingdi bundle
    Use this skill when a user wants evaluator data for judging answers about images, charts, infographics, screenshots, diagrams, or other visual inputs. Trigger it when people use plain requests like 'compare who answered the picture question better', 'score the chart explanation', 'rank several image-based answers', or 'make judge data for multimodal responses'. It is especially appropriate for scoring evaluation, pair comparison, or batch ranking in vision-language tasks where the judge must reason over both the visual content and the textual response. Example triggers include: 'evaluate answers about charts', 'judge which caption is better', 'rank image question responses', and 'make harder visual judge items with hallucinations'.
    0
    installs
  55. Multilingual Schema Grounding 2 · dingxingdi bundle
    Skill: multilingual schema grounding
    0
    installs
  56. Citation Grounded Report Synthesis 2 · dingxingdi bundle
    Use this skill when the user wants the agent to produce a sourced mini-report, a citation-rich answer, or a research-style output rather than a one-line fact. Trigger it for requests like “make it write a grounded report,” “the answer should cite sources,” or “test whether it can gather and synthesize evidence into a brief research note.” It is the right skill when success depends on both web evidence collection and disciplined, citation-backed synthesis.
    0
    installs
  57. Cross Lingual Ecosystem Navigation 2 · dingxingdi bundle
    Use this skill when the user wants the agent to leave the usual English web, search local-language sites, bridge naming variants, or answer an English prompt using non-English evidence. Trigger it for requests like “use Chinese websites,” “mix languages,” “make the key clue only appear on local sites,” or “test whether the agent can search beyond English pages.” It is appropriate when the challenge should come from cross-lingual retrieval, transliteration, and ecosystem-specific navigation rather than from English-only browsing.
    0
    installs
  58. Evidence Triage And Reconciliation 2 · dingxingdi bundle
    Use this skill when the user wants misleading pages, conflicting sources, false leads, or answer-like snippets that the agent must sort out instead of trusting blindly. Trigger it for requests like “make some pages look right but be wrong,” “force the model to compare conflicting evidence,” or “include misinformation and distractors.” It is especially useful when the target capability is deciding which evidence to trust, not just finding any evidence at all.
    0
    installs
  59. Element Complete Summarization 2 · dingxingdi bundle
    Skill: element-complete-summarization
    0
    installs
  60. Constraint Table Post Processing Reasoning 2 · dingxingdi bundle
    Use this skill when the user wants table-heavy questions, multi-constraint questions, ranking/filtering problems, or questions that require a small operation after reading the docs. Trigger it for requests like 'make it use a table', 'ask for the one that fits several conditions', 'force some calculation or formatting', or 'make the answer require post-processing'. This skill targets retrieval-plus-operation behavior.
    0
    installs
  61. Constraint Table Post Processing Reasoning 3 · dingxingdi bundle
    Use this skill when the user wants table-heavy questions, multi-constraint mathematical logic, SQL/programming tracing, or questions requiring post-processing formats. Trigger it for requests like 'make it use a table and text', 'trace the math logic of this textbook', 'ask for the one that fits several conditions', 'cross-reference the table with the article', or 'force an algorithmic calculation'. This skill targets complex operation execution after foundational text retrieval.
    0
    installs
  62. Descriptive And Statistical Analysis 2 · dingxingdi bundle
    Skill: descriptive and statistical analysis
    0
    installs
  63. Self Evolving Experience Reuse 2 · dingxingdi bundle
    Skill: self-evolving-experience-reuse
    0
    installs
  64. Objective Correctness Adjudication 2 · dingxingdi bundle
    Use this skill when a user wants evaluator data where the judge must decide which answer is actually correct, logically valid, or verifier-passing, especially in hard reasoning, math, coding, or knowledge tasks. Trigger it when people ask for things like 'make judge data that tests real correctness', 'compare correct and incorrect solutions', 'focus on logic before style', or 'build hard evaluator tasks that humans may misjudge'. Plain-language examples include: 'test whether the judge can pick the actually right answer', 'make response pairs with one correct and one wrong', 'evaluate reasoning quality objectively', and 'create judge data for hard math or coding answers'.
    0
    installs
  65. Multi Turn Revision Consistency 2 · dingxingdi bundle
    Skill: multi-turn-revision-consistency
    0
    installs
  66. Experiment Driven Model Improvement 2 · dingxingdi bundle
    Use this skill when the user wants data where the agent must improve an existing ML training setup by editing code and running experiments. Trigger it for requests like 'generate ML debugging and improvement tasks', 'make train.py optimization problems', 'create model tuning workflows with execution', or 'give me data where the agent has to read logs and improve accuracy'. Do not use it for open-ended Kaggle competition work from scratch or for paper-reproduction tasks.
    0
    installs
  67. Knowledge Update Tracking 2 · dingxingdi bundle
    Skill: knowledge-update-tracking
    0
    installs
  68. Temporal Ordering And Localization 2 · dingxingdi bundle
    Use this skill when the user wants 'when did this happen', 'which happened first', 'find the exact moment', or 'make the answer depend on ordering, timing, or the right segment of a long clip.' Trigger it for requests about timestamps, first-versus-second events, exact intervals, and questions that need watching the full sequence instead of spotting one frame.
    0
    installs
  69. Citation Grounded Factuality Checking 2 · dingxingdi bundle
    Use this skill when the user wants the agent to judge whether an answer is really backed by the document, not just whether it sounds correct. Trigger it for requests like "check if the citation really supports the claim," "make examples of sneaky hallucinations," "verify grounded answers sentence by sentence," or "test whether the model invents support from the source." It is ideal for legal research outputs, RAG answers, summaries, and any document-grounded generation setting where false support is as dangerous as an outright fabrication.
    0
    installs
  70. Safety Constrained Task Planning 2 · dingxingdi bundle
    Skill: safety-constrained-task-planning
    0
    installs
  71. Bias Resistant Human Aligned Judging 2 · dingxingdi bundle
    Use this skill when a user wants evaluator data that probes whether a judge is fair, stable, and aligned with human preferences instead of being swayed by position, length, self-style, names, or misleading cues. Trigger it when people say things like 'test evaluator bias', 'make judge robustness data', 'see if the evaluator changes when I swap the order', or 'build hard examples that look different on the surface but should get the same verdict'. Plain-language examples include: 'check if the judge is biased toward longer answers', 'test self-preference or name bias', 'make order-swap stability data', and 'create fairness-style evaluator tests'.
    0
    installs
  72. Parallel Subtask Scheduling 2 · dingxingdi bundle
    Skill: parallel-subtask-scheduling
    0
    installs
  73. Implicit Commonsense Consistency 2 · dingxingdi bundle
    Use this when the user wants planning data about plans that should 'make sense' even if the user did not spell out every rule. Trigger it for requests like 'make itinerary tasks where the plan should feel realistic', 'give me planning data that catches subtle nonsense', 'create questions where the agent must avoid repeating the same attraction or inventing fake flights', or 'make plans that need common sense instead of only exact keywords.'
    0
    installs
  74. Resource Lifecycle Orchestration 2 · dingxingdi bundle
    Use this when the user wants planning data about making something step by step, reusing tools correctly, cleaning items before reuse, or putting equipment back in order at the end. Trigger it for requests like 'make procedural planning tasks with cleanup', 'give me plans where tools have to be prepared before reuse', 'make maintenance tasks with strict order', or 'create plans where forgetting reset steps breaks the whole workflow.'
    0
    installs
  75. Contextual Entity Disambiguation 2 · dingxingdi bundle
    Skill: contextual entity disambiguation
    0
    installs
  76. State And Attribute Change Tracking 2 · dingxingdi bundle
    Use this skill when the user asks for questions about something changing over time, such as melting, darkening, opening, shrinking, turning on, or switching state. Trigger it for requests like 'make the clip look the same except for one change', 'ask what changed', 'did it open or close', or 'make the answer depend on the before-versus-after state.'
    0
    installs
  77. Long Form Breadth Depth Coherence 2 · dingxingdi bundle
    Skill: long-form-breadth-depth-coherence
    0
    installs
  78. Paper To Code Implementation Fidelity 2 · dingxingdi bundle
    Use this skill when the user wants data where the agent must read a paper-like specification and write the corresponding code from scratch or near-scratch. Trigger it for requests like 'generate paper-to-code tasks', 'make research implementation data', 'create long-horizon coding tasks from papers', or 'give me ML R&D examples where the agent must turn methods into code'. Do not use it for ordinary repository issue fixing or for later-stage reproduction verification alone.
    0
    installs
  79. Persona Consistent Roleplay 2 · dingxingdi bundle
    Skill: persona-consistent-roleplay
    0
    installs
  80. Open World Object Goal Navigation 2 · dingxingdi bundle
    Skill: open-world-object-goal-navigation
    0
    installs
  81. Long Horizon Itinerary Sequencing 2 · dingxingdi bundle
    Use this when the user wants planning data about multi-day schedules, several cities, or chains of decisions that must stay consistent from one day to the next. Trigger it for requests like 'make week-long itinerary tasks', 'give me travel plans that span multiple stops', 'create schedule planning data where each day depends on the last one', or 'make long plans where an early routing choice affects later days.'
    0
    installs
  82. Multi Document Integrative Reasoning QA 2 · dingxingdi bundle
    Use this skill when the user wants questions that require connecting information across several files or several far-apart sections of a long workspace. Trigger it for requests like "make it require combining documents," "ask something that needs comparing several reports," "force the agent to connect the dots," or "the answer should depend on multiple sources, not one page." It is the right skill for benchmark-style document bundles, deep-research corpora, meeting materials, legal case collections, or paper sets where skipping one document should break the answer.
    0
    installs
  83. Visualization Grounded Insight Reporting 2 · dingxingdi bundle
    Skill: visualization-grounded insight reporting
    0
    installs
  84. Agent Trajectory Safety Risk Diagnosis 2 · dingxingdi bundle
    Use this skill when a user wants evaluator data for judging whether an agent’s actions were safe or unsafe based on its interaction record with tools or environments. Trigger it when ordinary users say things like 'check if the agent did something dangerous', 'judge whether the action trail is risky', 'make safety-monitor data from agent logs', or 'find cases where the agent looked helpful but caused harm'. Plain-language examples include: 'evaluate whether the trajectory is safe', 'make judge data for risky tool use', 'analyze unsafe multi-step agent actions', and 'test if the evaluator can spot hidden safety problems in logs'.
    0
    installs
  85. Live Web Multimodal Information Seeking 2 · dingxingdi bundle
    Use this skill when the user wants questions that can only be solved by actually using the web page, video, image, map, tour, or other live interactive content. Trigger it for requests like “make it find the answer from the real page,” “don’t let search snippets be enough,” “use a video or interactive asset,” or “force browsing plus media understanding.” This skill is for GUI-browser tasks where the answer is buried in a specified source and must be extracted through real rendered interaction.
    0
    installs
  86. Subjective Qualitative Response Scoring 2 · dingxingdi bundle
    Use this skill when a user wants single-response judge data that scores answer quality on dimensions ordinary people describe as 'clear', 'concise', 'complete', 'formal enough', 'follows the instruction', or 'sounds well written'. Trigger it when the task is not just right-versus-wrong, but rating one response against a qualitative standard. Example plain-language triggers include: 'score how well the answer follows the prompt', 'make evaluator data for concise but complete answers', 'judge whether the response is clear and on-topic', and 'test if the model rambles or misses the format request'.
    0
    installs
  87. Dependency Aware Tool Chaining 2 · dingxingdi bundle
    Skill: dependency-aware-tool-chaining
    0
    installs
  88. Verifiable Translation Rule Compliance 2 · dingxingdi bundle
    Use this skill when the user wants translation data with hard rules that can be checked exactly, such as "keep the URL as-is," "change dates to DD/MM/YYYY," "convert dollars to pesos," or "put the original museum name in parentheses." Trigger it for requests like "test translation plus formatting rules," "make it follow localization instructions," or "the output must obey several exact constraints, not just be a good translation."
    0
    installs
  89. Discourse Consistent Document Translation 2 · dingxingdi bundle
    Use this skill when the user wants translation tasks that break sentence-by-sentence systems and require document context. Trigger it for requests like "make the translation depend on earlier sentences," "test pronoun resolution across the paragraph," "preserve who is speaking," or "keep the wording and style consistent across the whole passage." It is especially useful for literary paragraphs, dialogue scenes, and technical documents where the wrong pronoun, tense anchor, or repeated term can destroy coherence.
    0
    installs
  90. Episodic And Multi Hop Memory Reasoning 2 · dingxingdi bundle
    Use this skill when the user wants questions like 'last time this happened', 'after they went there, what did they do next', 'combine clues from different parts', or 'connect the dots across a very long video.' Trigger it when the answer cannot be read from one clip and instead depends on stitching together multiple past moments.
    0
    installs
  91. Faithful Fine Grained Video Description 2 · dingxingdi bundle
    Use this skill when the user wants 'describe the video in detail', 'cover all important actions', 'write a rich but accurate description', or 'make captioning depend on many events instead of one obvious action.' Trigger it for open-ended generative data where the challenge is not choosing an option but producing a faithful, event-complete description.
    0
    installs
  92. Schema Bound Document Attribute Extraction 2 · dingxingdi bundle
    Use this skill when the user wants structured labels or fields extracted from messy document text. Trigger it for requests like "label the sentence under this schema," "pull the key attributes from the note," "classify the document with strict categories," or "extract fields only when the evidence truly matches the guideline." It is appropriate for clinical notes, legal clauses, financial agreements, contracts, and any document workflow that converts prose into a controlled schema.
    0
    installs
  93. Explicit Hard Constraint Satisfaction 2 · dingxingdi bundle
    Use this when the user wants planning data about strict requirements like budget caps, room type, no pets, no flight, cuisine preferences, or other must-have conditions. Trigger it for requests like 'make planning tasks with firm user requirements', 'give me itinerary problems with budget and hotel rules', 'create plans that must obey no-flight or no-driving constraints', or 'make data where the plan must fit several stated preferences at once.'
    0
    installs
  94. Tool Mediated Information Acquisition 2 · dingxingdi bundle
    Use this when the user wants planning data about looking things up before deciding, choosing the right search tool, or gathering scattered facts into one plan. Trigger it for requests like 'make planning tasks where the agent has to search first', 'give me itinerary problems that require checking flights and hotels', 'create planning tasks with missing information', or 'make questions where the plan depends on tool use instead of guessing.'
    0
    installs
  95. Multi Turn Instruction Following 2 · dingxingdi bundle
    Skill: multi-turn-instruction-following
    0
    installs
  96. Landmark Grounded Semantic Exploration 2 · dingxingdi bundle
    Skill: landmark-grounded-semantic-exploration
    0
    installs
  97. Hierarchical Rubric Based Artifact Grading 2 · dingxingdi bundle
    Use this skill when a user wants evaluator data for grading complex artifacts against a rubric, such as codebases, research replications, long reports, logs, or multi-file submissions. Trigger it when the request sounds like 'break the task into rubric checks', 'grade the artifact against detailed criteria', 'evaluate partial progress', or 'make judge data for complex outputs that cannot be scored by one simple metric'. Plain-language examples include: 'judge whether the project meets the checklist', 'score a submission using leaf-node criteria', 'evaluate artifact quality step by step', and 'build judge data for rubric-based grading'.
    0
    installs
  98. Causal And Counterfactual Video Reasoning 2 · dingxingdi bundle
    Use this skill when the user asks for 'why did that happen', 'what caused it', 'what would happen if this were removed', or 'make the answer depend on reasoning rather than direct observation.' Trigger it for clips where the visible events must be interpreted through motive, mechanism, or hypothetical intervention.
    0
    installs
  99. Interactive Information Gathering 2 · dingxingdi bundle
    Skill: interactive-information-gathering
    0
    installs
  100. Focus Controlled Element Aware Summarization 2 · dingxingdi bundle
    Use this skill when the user wants a summary with a clear lens or checklist rather than a generic paragraph. Trigger it for requests like "cover the key who/when/what/result," "summarize only the parts about X," "make the summary follow a fixed frame," or "focus on the specific question, not the whole document equally." This is especially useful for news, meetings, multi-news bundles, reviews, or technical documents where the task is selective coverage under an explicit focus constraint.
    0
    installs