Skip to content

Instantly share code, notes, and snippets.

@kaapstorm
Created July 17, 2026 15:02
Show Gist options
  • Select an option

  • Save kaapstorm/b612e270e34906a392de8b01c7d792f8 to your computer and use it in GitHub Desktop.

Select an option

Save kaapstorm/b612e270e34906a392de8b01c7d792f8 to your computer and use it in GitHub Desktop.
Using Ollama + OpenCode + Qwen3.5-9B

Using Ollama + OpenCode + Qwen3.5-9B

Wat?

Ollama is a command line utility for installing and running language models.

OpenCode is an open-source equivalent of Claude Code. It is an agent harness recommended for running local models.

Getting Started

  1. Install Ollama and OpenCode.

    $ curl -fsSL https://ollama.com/install.sh | sh
    $ curl -fsSL https://opencode.ai/install | bash

    (Optional) Avoid another directory on your path by creating a symlink to OpenCode.

    $ cd ~/.local/bin/
    $ ln -s $HOME/.opencode/bin/opencode
  2. Install Qwen 3.6.

    1. Start Ollama.

      $ ollama
    2. Under "Recommended", select "qwen3.6 (install)".

    3. Wait for the 23 GB download to complete.

  3. Investigate your new model.

    >>> /show info
      Model
        architecture        qwen35moe    
        parameters          36.0B        
        context length      262144       
        embedding length    2048         
        quantization        Q4_K_M       
    
      Capabilities
        completion    
        vision        
        tools         
        thinking      
    
      Parameters
        min_p               0       
        presence_penalty    1.5     
        repeat_penalty      1       
        temperature         1       
        top_k               20      
        top_p               0.95    
    
      License
        Apache License               
        Version 2.0, January 2004    
        ...                          
    
    >>> /show system
    No system message was specified for this model.
    
    >>> /show template
    {{ .Prompt }}
    
  4. Say something.

    >>> By what name would you like to be addressed?
    Thinking...
    Here's a thinking process:
    
    1.  **Analyze User Input:**
       - User asks: "By what name would you like to be addressed?"
       - This is a direct question about my preferred name/how I'd like to be addressed.
    
    2.  **Identify Key Constraints/Context:**
       - I am an AI assistant developed by Alibaba Group's Tongyi Lab.
       - My official name is Qwen (or Tongyi Qianwen).
       - I should respond in a helpful, respectful, and clear manner.
       - I should acknowledge my identity while leaving room for user preference if they want to address me differently.
    
    3.  **Formulate Response Strategy:**
       - Acknowledge the question politely.
       - State my official name clearly (Qwen / Tongyi Qianwen).
       - Note that I'm flexible and happy to be addressed however the user prefers.
       - Keep it concise and friendly.
    
    4.  **Draft Response (Mental Refinement):**
       You can call me Qwen (or Tongyi Qianwen), as that's my official name. But I'm perfectly happy to be addressed however you'd 
    prefer—just let me know what works best for you!
    
    5.  **Check Against Guidelines:**
       - Accurate to identity? Yes, Qwen/Tongyi Qianwen.
       - Clear and concise? Yes.
       - Matches tone? Friendly and helpful.
       - No extra fluff? Yes.
    
    6.  **Final Output Generation:** (matches the drafted response)✅
       "You can call me Qwen (or Tongyi Qianwen), which is my official name. But I’m perfectly happy to be addressed however you’d 
    prefer—just let me know what works best for you!" 
    
       This is accurate, flexible, and aligns with my design. Proceed. 
       Output matches response.✅
    ...done thinking.
    
    You can call me **Qwen** (or Tongyi Qianwen), which is my official name. But I'm perfectly happy to be addressed however you'd 
    prefer—just let me know what works best for you!
    

But it's sooo slooowww

For non-coding tasks, like managing a to-do list, a smaller model will be capable, and faster. Other models to try:

  • Qwen 3.5 9B: "Best general model on 8GB | Fits on RTX 3060/4060, 8GB"
  • Qwen 3.6 27B: "Best open coder on a single 24GB card | RTX 3090/4090, 24GB"
$ ollama pull qwen3.5:9b

Using with OpenCode

Running a model in Ollama just gives you a chat interface. To run an agent, you need a harness, like Claude Code. OpenCode and Pi are open-source harnesses that can run local (or cloud) models. Pi is a bare-bones harness without guardrails. OpenCode is an equivalent to Claude Code.

Ollama defaults to a 4096-token context window.

OpenCode needs at least 16k–64k to function reliably. In practice 32k often works for smaller repos. Go higher if you can afford the VRAM (GPU memory).

  1. Create a model variant with a larger context window.

    $ ollama run qwen3.5:9b
    >>> /set parameter num_ctx 32768
    Set parameter 'num_ctx' to '32768'
    >>> /save qwen3.5:9b32k
    Created new model 'qwen3.5:9b32k'
    >>> /bye
  2. Edit ~/.config/opencode/opencode.jsonc. (It should already exist. Older documentation also refers to ~/.config/opencode/config.json or ~/.config/opencode/opencode.json.)

    {
      "$schema": "https://opencode.ai/config.json",
      "provider": {
        "ollama": {
          "npm": "@ai-sdk/openai-compatible",
          "name": "Ollama (local)",
          "options": {
            "baseURL": "http://localhost:11434/v1"
          },
          "models": {
            "qwen3.5:9b32k": {
              "name": "qwen3.5:9b32k"
            }
          }
        }
      }
    }
  3. You are ready:

    $ opencode --model ollama/qwen3.5:9b32k
                                       ▄
      █▀▀█ █▀▀█ █▀▀█ █▀▀▄ █▀▀▀ █▀▀█ █▀▀█ █▀▀█
      █  █ █  █ █▀▀▀ █  █ █    █  █ █  █ █▀▀▀
      ▀▀▀▀ █▀▀▀ ▀▀▀▀ ▀  ▀ ▀▀▀▀ ▀▀▀▀ ▀▀▀▀ ▀▀▀▀

    Ask it to do something.

A note regarding LSPs

OpenCode offers LSP integration. To enable LSPs, add "lsp": true to your opencode.jsonc:

{
  "$schema": "https://opencode.ai/config.json",
  "lsp": true,
  "provider": {
    // ...
  }
}

For Python, the built-in pyright LSP requires pyright to be installed as a dependency in your project.

LSP integration can help the agent find and fix issues by providing diagnostics, but it's not always a net positive. Language servers can use significant memory, and slow down agent workflows. For many projects it's better just to have the agent run your lint/typecheck commands directly.

Why use a local model?

  • You don't have to pay an AI company for usage.

  • It's private. If you want to manage your to-do list, and if you have information in there that you don't want to upload to a cloud environment that isn't owned by Google or managed by Dimagi, this is a good fit.

More info

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment