Skip to content

Instantly share code, notes, and snippets.

@piotrekkaminski
Created August 14, 2026 15:51
Show Gist options
  • Select an option

  • Save piotrekkaminski/d3312b94d26caf57f50b2ce36578ebc7 to your computer and use it in GitHub Desktop.

Select an option

Save piotrekkaminski/d3312b94d26caf57f50b2ce36578ebc7 to your computer and use it in GitHub Desktop.
creative — a Claude/agent skill implementing Verbalized Sampling (arXiv:2510.01171). Generates 10 substantially different long-tail responses with explicit probability estimates.
name creative
description Generate 10 substantially different, long-tail responses to a question or task, each with an estimated probability under 0.10. Use when the user asks for creative thinking, a creative approach, "be creative", "think creatively", "give me creative options", "what are the unconventional angles", "sample the long tail", or otherwise signals that the obvious answer is not what they want. Applies to product and strategy ideation, reframing stuck problems, and finding fresh angles for writing. Do NOT use for routine execution, factual lookups, or when the user wants the single best conventional answer.

Creative

Sample the long tail. The user already knows the obvious answer. This skill exists to surface the answers that ordinary prompting would bury.

Based on Verbalized Sampling. See Attribution at the bottom.

When this runs

The user asks for creative thinking, a creative approach, unconventional angles, or says something like "be creative", "think creatively", "what else could this be", "sample the long tail". It covers three kinds of work with the same method:

  • Product and strategy ideation (features, experiments, positioning, pitches)
  • Reframing a stuck problem (the framing itself may be wrong)
  • Writing and narrative (hook, structure, metaphor, angle)

Do not run it for routine execution, factual lookups, or when the user wants the single best conventional answer.

The method

Generate 10 substantially different responses to the query. Each response must contain:

  • the response text
  • probability: an estimated probability from 0.0 to 1.0 relative to the full distribution of plausible responses

Sample from the long tail of the distribution. Prefer valid responses that would normally receive relatively little probability mass under ordinary prompting. Aim for each candidate to have estimated probability below 0.10.

Do not make responses strange merely for novelty. They must remain logically sound, relevant, and useful.

Explore different:

  • assumptions
  • conceptual frameworks
  • solution strategies
  • interpretations
  • edge cases
  • contrarian but defensible possibilities

What "substantially different" means

Two responses are not different because they use different words. They are different when they rest on a different assumption, use a different framework, or would lead to a different next action. If two candidates would produce the same work if acted on, one of them is wasted. Replace it.

Common failure: ten variations of the same idea at different levels of detail. Check for this before answering.

Calibrating the probability

The number is an honest estimate of how often this response would appear if the question were asked many times under normal conditions. Do not inflate it to look confident, and do not deflate it to look adventurous.

If a candidate lands above 0.10, it is the obvious answer wearing a costume. Either cut it or keep exactly one as a baseline and label it as such, so the distance of the others is visible.

Output format

**1. <short handle>** - p ~ 0.0X
<the response>

**2. <short handle>** - p ~ 0.0X
<the response>
...

Handles are short so the user can point at one and say "expand 4".

No preamble. No summary of what you are about to do. Start at candidate 1.

After the list

One line only: offer to expand any candidate, or to converge the set into a shortlist with reasoning. Do not converge unless asked. The point of the list is that the sorting happens after the user has seen the range.

Attribution

The core technique is Verbalized Sampling, from Zhang et al., Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity (arXiv:2510.01171).

The paper's argument in short: alignment training collapses a model's output distribution around a narrow set of typical, familiar answers. The root cause is not the optimization algorithm but typicality bias in human preference data, since annotators systematically prefer responses that sound conventional. The long tail of the base model's distribution is still there, just suppressed. Asking the model to verbalize a distribution over multiple candidates, with explicit probabilities and a bias toward the low-probability region, recovers much of it. The paper reports 1.6x to 2.1x semantic diversity gains over direct prompting and around 66.8% of base-model diversity recovered, with quality, factual accuracy and safety behaviour maintained.

The prompt template this skill is built around comes from Brian Roemmele's write-up of the paper: https://x.com/brianroemmele/status/2088069286850179288

Everything under "What substantially different means", "Calibrating the probability", "Output format" and "After the list" is added scaffolding, not from the paper. It exists because a bare template drifts once it runs inside a longer agent session.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment