Skip to content

Instantly share code, notes, and snippets.

@opparco
Created May 6, 2026 10:11
Show Gist options
  • Select an option

  • Save opparco/755173410fd5ae3ef42e42880c2769ad to your computer and use it in GitHub Desktop.

Select an option

Save opparco/755173410fd5ae3ef42e42880c2769ad to your computer and use it in GitHub Desktop.
Verification of Artist Tag Case Sensitivity in Anima

Verification of Artist Tag Case Sensitivity in Anima

Overview

This report verifies, at the codebase level, whether artist tags in circlestone-labs/Anima are treated as case-insensitive, meaning camelcase and lowercase inputs would be processed as the same prompt text.

Target strings:

  • @onikobe rin
  • @Onikobe rin

The conclusion is that they are not treated as identical in this codebase. They diverge at the tokenizer stage, which means the resulting text embeddings can also differ.

Implementation Evidence

1. Anima uses the Qwen tokenizer and T5 tokenizer directly

In backend/diffusion_engine/anima.py, the Anima text processing engine receives tokenizer and tokenizer_2 directly.

Relevant snippet:

clip = CLIP(
    model_dict={"qwen3_06b": huggingface_components["text_encoder"]},
    tokenizer_dict={
        "qwen3_06b": huggingface_components["tokenizer"],
        "t5xxl": huggingface_components["tokenizer_2"],
    },
)

self.text_processing_engine_anima = AnimaTextProcessingEngine(
    text_encoder=clip.cond_stage_model.qwen3_06b,
    qwen_tokenizer=clip.tokenizer.qwen3_06b,
    t5_tokenizer=clip.tokenizer.t5xxl,
)

2. No case normalization is applied when tokenizers are loaded

In backend/loader.py, tokenizer components are loaded with from_pretrained() and returned as-is. There is no preprocessing step such as lower() or casefold().

Relevant snippet:

if component_name.startswith("tokenizer"):
    cls = getattr(importlib.import_module(lib_name), cls_name)
    comp = cls.from_pretrained(os.path.join(repo_path, component_name))
    comp._eventual_warn_about_too_long_sequence = lambda *args, **kwargs: None
    return comp

3. AnimaTextProcessingEngine passes the input text directly to the tokenizers

In backend/text_processing/anima_engine.py, tokenize() sends the input text directly to both the Qwen tokenizer and the T5 tokenizer. No case normalization is performed here either.

Relevant snippet:

def tokenize(self, texts):
    return (
        self.qwen_tokenizer(texts, truncation=False, add_special_tokens=False)["input_ids"],
        self.t5_tokenizer(texts, truncation=False, add_special_tokens=False)["input_ids"],
    )

Tokenizer Measurement Results

The bundled local Anima tokenizers were used directly to compare the token IDs for the two target strings.

Qwen tokenizer

@onikobe rin  -> [31, 263, 1579, 15422, 53721]
@Onikobe rin  -> [31, 1925, 1579, 15422, 53721]

False

T5 tokenizer_2

@onikobe rin  -> [3320, 106, 12027, 346, 3, 52, 77]
@Onikobe rin  -> [3320, 7638, 12027, 346, 3, 52, 77]

False

Because the token IDs differ in both tokenizers, @onikobe rin and @Onikobe rin are processed as different inputs.

Conclusion

In this codebase, artist tag handling for circlestone-labs/Anima is not case-insensitive.

  • There is no lowercase or casefold normalization in the relevant code path
  • The Qwen tokenizer produces different token IDs
  • tokenizer_2 also produces different token IDs

Therefore, camelcase and lowercase artist tags can produce different embeddings, so the report that they lead to different outputs is consistent with the implementation in this codebase.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment