Skip to content

Instantly share code, notes, and snippets.

@cidrblock
Last active April 17, 2026 11:54
Show Gist options
  • Select an option

  • Save cidrblock/eaa433824950742554bee78d217e3772 to your computer and use it in GitHub Desktop.

Select an option

Save cidrblock/eaa433824950742554bee78d217e3772 to your computer and use it in GitHub Desktop.
Selenium → WebDriverIO migration analysis for vscode-ansible

Selenium → WebDriverIO migration analysis for vscode-ansible

Selenium to WebDriverIO (WDIO) Migration

Executive Summary

The Selenium UI test suite ran Python pytest against a containerized code-server instance via Selenium Grid. This setup had a 71% CI failure rate, primarily due to infrastructure fragility (container health checks, code-server startup timing, multi-service container coordination). It also tested code-server, not VS Code -- so passing tests did not guarantee real-world correctness.

We replaced it with WebDriverIO (WDIO) using wdio-vscode-service, which tests against real VS Code (Electron), eliminates all container infrastructure, and uses the same language (TypeScript) as the extension.

At the time of replacement, all 29 Selenium tests on main were skipped — zero tests actually executed in CI. WDIO replaces them with 30 passing tests across 7 spec files, including mock-based Lightspeed coverage that never existed in CI.

PR: ansible/vscode-ansible#2746


CI Failure History

The Selenium container (docker-compose.yml) was introduced on 2026-02-10 in PR #2560. The following table shows weekly CI pass/fail rates on the main branch from that point forward.

Data source: gh api repos/ansible/vscode-ansible/actions/workflows/ci/runs (300 completed runs sampled, Feb 10 -- Apr 16, 2026).

Week Runs Pass Fail Cancel Fail Rate Primary Failure
W09 (Feb 24) 19 18 1 0 5% builder-image (container build)
W10 (Mar 3) 47 43 3 1 6% builder-image
W11 (Mar 10) 47 37 4 6 9% mixed
W12 (Mar 17) 51 43 8 0 16% builder-image + task ui (selenium)
W13 (Mar 24) 40 25 15 0 38% task ui (selenium) + WSL runner
W14 (Mar 31) 42 29 13 0 31% task ui (selenium)
W15 (Apr 7) 33 0 33 0 100% task ui (selenium) -- every run
W16 (Apr 14) 21 0 21 0 100% task ui (selenium) -- every run

Key observations:

  • Selenium tests were stable for ~3 weeks (W09-W11), then degraded steadily as the container image DOM diverged from the XPath selectors.
  • From W15 onward, every single CI run on main failed due to Selenium, making the main branch permanently red.
  • The Selenium failure rate over the full period: 98 failures out of 300 runs (33%), with the rate accelerating to 100% in the final two weeks.
  • Non-Selenium failures (e2e timeout, WSL runner, container build) accounted for a minority of total failures.

Selenium Test Status on main (at time of removal)

All 29 Selenium tests on main were @pytest.mark.skip — the CI step still ran (built the container, invoked pytest) but executed zero actual tests:

File Tests Status Skip Reason
test_00_commands.py 2 All skipped Flaky on CI - extension activation timing
test_01_dev_webviews.py 2 All skipped Flaky on CI - extension activation timing
test_02_welcome_webviews.py 3 All skipped Flaky on CI - extension activation timing
test_03_llm_provider_webview.py 4 All skipped Flaky on CI - extension activation timing
test_50_mcp_server.py 4 All skipped Flaky on CI - extension activation timing
test_80_lightspeed.py 5 All skipped AAP-67210
test_81_login.py 8 All skipped AAP-67210
test_82_lightspeed_trial.py 1 All skipped AAP-67210
Total 29 0 running
  • 15 tests skipped for "Flaky on CI - extension activation timing issues" — the core UI tests (commands, webviews, MCP).
  • 14 tests skipped for AAP-67210 — all Lightspeed and SSO tests.

AAP-67210 Coverage

AAP-67210 tracks 14 Lightspeed/SSO tests that were disabled in Selenium. WDIO restores coverage for 5 using a mock server, and adds 2 new tests:

AAP-67210 Selenium Test WDIO Status WDIO Test
test_vscode_widget Covered should have Lightspeed commands registered
test_vscode_playbook_explanation Covered should open playbook explanation webview
test_vscode_playbook_generation Covered should open playbook generation webview
test_vscode_role_generation Covered should open role generation webview
test_vscode_lightspeed_explorer Partial Explorer view covered via sidebar tests
test_vscode_trial_button Deferred Excluded from initial PR
test_unsubed_login N/A SSO flow — tests RH infra, not extension
test_unsubed_admin_login N/A SSO flow
test_no_wca_user_login N/A SSO flow
test_no_wca_admin_login N/A SSO flow
test_login_page N/A SSO flow
test_sso_auth_flow N/A SSO flow
test_admin_portal_error N/A SSO flow
test_vscode_rhsso_auth_flow N/A SSO flow

The 8 SSO login tests exercise Red Hat SSO infrastructure (external OAuth pages, live credentials, running Lightspeed backend), not extension code. They were never functional in CI regardless of test framework.

Net-new tests not in AAP-67210: should handle API errors gracefully and should handle server timeout gracefully — error/timeout resilience that was never tested.


Before: Selenium Architecture

Host (GitHub Actions runner)
├── Python pytest
│   └── Selenium RemoteWebDriver → http://localhost:4444/wd/hub
│
└── Podman container (ghcr.io/ansible/selenium-adt:main)
    ├── Selenium Grid (:4444)
    ├── Firefox/Chrome browser
    ├── code-server (:8080) ← NOT VS Code
    │   └── Ansible extension (from mounted VSIX)
    └── VNC (:5999)

Problems

Issue Impact
Health check only validated Selenium Grid, not code-server 71% CI failure rate
code-server != VS Code (DOM differences, API gaps) Tests may pass but not reflect real user experience
Three services in one container Complex startup coordination
Python tests, TypeScript extension No shared types or utilities
Raw XPath selectors + manual iframe traversal Brittle, breaks on UI changes
External container image (selenium-adt) Build/publish pipeline outside this repo
~10-20 min runtime on CI Blocks release pipeline

After: WDIO Architecture

Host (GitHub Actions runner or local workstation)
├── WDIO test runner (TypeScript)
│   └── wdio-vscode-service
│       ├── Downloads matching VS Code + Chromedriver
│       ├── Launches VS Code Electron with extension
│       └── Provides page objects (Workbench, EditorView, etc.)
│
└── VS Code Electron (real desktop app)
    └── Ansible extension (extensionDevelopmentPath)

No containers. No Selenium Grid. No code-server. No Python.


What Was Removed

Item Description
test/ui/ (entire directory) Python pytest tests, fixtures, XPath helpers (700+ lines)
docker-compose.yml Selenium/code-server container orchestration
podman-compose (pyproject.toml) Python dependency for container management
selenium (pyproject.toml) Python dependency for WebDriver
Container health checks Startup coordination for 3 in-container services
VNC debugging Replaced by native VS Code and screenshots

What Was Added

File Purpose Lines
wdio.conf.ts WDIO configuration (VS Code service, Mocha, specs) 57
test/wdio/smoke.spec.ts Extension activation, activity bar, commands, language detection 99
test/wdio/terminal.spec.ts adt --version availability, terminal API 43
test/wdio/commands.spec.ts ansible.create-empty-playbook command 71
test/wdio/webviews.spec.ts Devfile, devcontainer, welcome page, LLM settings 219
test/wdio/mcp.spec.ts MCP enable/disable via commands and settings 174
test/wdio/sidebar.spec.ts ADT sidebar, welcome page link, settings propagation 90
test/wdio/lightspeed.spec.ts Mock-based Lightspeed tests (new coverage) 303
test/wdio/mock-server.ts Express mock server for Lightspeed API 190
test/wdio/fixtures/ Playbook files and VS Code settings for tests —
scripts/install-test-extensions.mjs Downloads VS Code + installs dependency extensions 43
test/wdio/tsconfig.json TypeScript config for WDIO specs —

Extension source changes (minimal)

File Change
src/extension.ts Added ansible.lightspeed.mockSession command for test session injection
src/features/lightspeed/lightSpeedOAuthProvider.ts Added setMockSession() method to write fake auth sessions

The mockSession command is not contributed in package.json, has no UI surface, and only writes to the extension's local secret store. It exists solely to let WDIO tests bypass SSO without live credentials.


Test Coverage Comparison

Ported from Selenium (16 tests)

# Selenium test WDIO spec Approach
1 test_terminal terminal.spec.ts execSync("adt --version") + terminal API
2 test_create_empty_playbook commands.spec.ts executeWorkbench() + editor content
3 test_devfile_webview webviews.spec.ts Webview context switch + content polling
4 test_devcontainer_webview webviews.spec.ts Same pattern
5 test_header_and_subtitle webviews.spec.ts Welcome page header assertion
6 test_header_subtitle webviews.spec.ts Welcome page subtitle assertion
7 test_mcp_section webviews.spec.ts MCP section in welcome page
8-11 test_llm_provider_* (x4) webviews.spec.ts LLM settings webview (list, edit, connect)
12-15 test_mcp_server_* (x4) mcp.spec.ts Enable/disable via commands and settings
16 test_sidebar_nav sidebar.spec.ts Activity bar + sidebar view

Net-new coverage (14 tests, never existed in CI)

# Test Spec file What it validates Why Selenium couldn't
17 VS Code launch smoke.spec.ts Session starts successfully No validation of this
18 Activity bar icon smoke.spec.ts Ansible icon visible in activity bar XPath-dependent
19 Extension activation smoke.spec.ts isActive === true via VS Code API No vscode API access
20 Command registration smoke.spec.ts ansible.* commands exist Couldn't enumerate
21 Language detection smoke.spec.ts .ansible.yml → languageId "ansible" No API; indirect
22 Welcome page link sidebar.spec.ts Sidebar opens welcome page Never tested
23 Settings propagation sidebar.spec.ts Runtime config change round-trip Pre-baked only
24 Terminal API terminal.spec.ts Create and dispose terminal Only scraped DOM
25 Lightspeed commands lightspeed.spec.ts Commands registered after mock login Required live SSO
26 Playbook explanation lightspeed.spec.ts Mock API → explanation webview Required live SSO
27 Playbook generation lightspeed.spec.ts Mock API → generation webview Required live SSO
28 Role generation lightspeed.spec.ts Mock API → role webview Required live SSO
29 API error handling lightspeed.spec.ts 500 → graceful degradation Never tested
30 Timeout handling lightspeed.spec.ts Slow server → timeout behavior Never tested

Total: 30 tests vs 16 active Selenium tests (88% increase).

The Lightspeed tests (#25-30) are the most significant gain: they test end-to-end UI flows that were permanently skipped in Selenium because they required live Red Hat SSO credentials and a running Lightspeed backend. WDIO tests use a local Express mock server and injected authentication sessions instead.


Key Implementation Details

Mock Lightspeed server

A single Express server (test/wdio/mock-server.ts) stubs all Lightspeed API endpoints with configurable canned responses:

POST /api/v0/ai/completions/      → inline suggestions
POST /api/v0/ai/explanations/     → playbook explanations
POST /api/v0/ai/generations/      → playbook/role generation
GET  /api/v0/me/                  → user info

Individual tests can override any endpoint via setResponse(endpoint, status, body). Overrides are cleared between tests with resetResponses().

Authentication mocking

Instead of real SSO, tests call a ansible.lightspeed.mockSession command that injects a synthetic auth session into the extension's secret store:

await browser.executeWorkbench(async (vscode) => {
  await vscode.commands.executeCommand(
    "ansible.lightspeed.mockSession",
    { accessToken: "mock", accountId: "test", accountLabel: "Test" },
  );
});

browser.executeWorkbench()

WDIO's executeWorkbench runs code inside the VS Code process with full vscode API access. This replaces XPath/CSS selectors for assertions that aren't about the DOM:

  • Check extension activation state
  • Read/write configuration values
  • Execute commands by ID
  • Open files and check language IDs
  • Query editor content

Webview testing

WDIO's switchToFrame() handles VS Code webview iframe traversal automatically. A polling helper waitForWebviewText() retries until Vue hydration completes, avoiding the race conditions that plagued the Selenium iframe navigation.


CI Integration

Before

- name: Configure podman
  run: ...
- name: Run UI tests (Selenium)
  run: task ui  # podman-compose up → pytest → podman-compose down

After

- name: Run UI tests (WDIO)
  run: task wdio  # pretest:wdio → xvfb-run npx wdio

WDIO runs on the same Linux runner as unit and e2e tests. No containers, no podman, no extra runner. Uses xvfb-run for headless display (same as the existing e2e tests).

Running locally

task wdio                        # full suite
npx wdio --spec smoke            # single spec
npx wdio --spec lightspeed       # lightspeed tests only

Requires xvfb-run on Linux (Wayland users: sudo dnf install xorg-x11-server-Xvfb).


Overlap with E2E Tests

The test/e2e/ suite (@vscode/test-cli + Mocha) tests the extension API layer inside VS Code Electron:

E2E Coverage UI Coverage (WDIO) Overlap
Hover providers (with/without EE) — None
ansible-lint diagnostics — None
YAML validation diagnostics — None
Outside-workspace LS — None
MCP activation (API-level) MCP enable/disable (UI-level) Complementary
— Terminal, commands None
— Webviews (devfile, welcome, LLM) None
— Lightspeed flows None

The two suites test different layers with minimal overlap. Both are kept. WDIO replaced only the Selenium UI suite.


What's Excluded (Future Work)

Item Why
Full SSO login flow (external RH SSO page) Tests Red Hat infrastructure, not the extension
Admin portal error test Same
Trial button test Deferred from initial PR
Inline suggestion ghost text Technically possible but tricky; deferred

The SSO/OAuth tests were skipped in Selenium (AAP-67210) and remain out of scope. They test Red Hat infrastructure, not extension behavior.


Side-by-Side: Selenium vs WDIO

Example 1: Terminal test — verify adt --version

Selenium (Python, 35 lines)

The test opens the command palette via XPath, sends a terminal command through sendSequence, then scrapes the terminal DOM to find output:

def test_terminal(browser_setup, screenshot_on_fail, close_editors):
    driver, _ = browser_setup
    ensure_vscode_ready(driver)                                     # 120s wait
    vscode_run_command(driver, ">workbench.action.terminal.new")    # F1 → XPath
    vscode_run_command(driver, ">workbench.action.terminal.focus")
    time.sleep(3)  # arbitrary sleep
    vscode_run_command(
        driver, ">workbench.action.terminal.sendSequence",
        "adt --version\\n"
    )
    output = None
    def check_output():
        nonlocal output
        output = driver.find_element(
            by="xpath", value="//div[@class='terminal-xterm-host']"
        )
        return "tox-ansible" in output.text
    errors = [NoSuchElementException, ElementNotInteractableException]
    wait = WebDriverWait(driver, timeout=10, poll_frequency=0.5,
                         ignored_exceptions=errors)
    wait.until(lambda _: check_output())
    text = output.text
    missing = [pkg for pkg in EXPECTED_ADT_PACKAGES if pkg not in text]
    assert not missing

Hidden cost: vscode_run_command is a 50-line helper that clicks the command center via XPath, retries up to 6 times with time.sleep(1), and manually sends keystrokes to the input box.

WDIO (TypeScript, 10 lines)

Runs adt --version directly in the test process — no DOM scraping:

it("should have adt available", async function () {
  this.timeout(30_000);
  const output = execSync("adt --version", {
    encoding: "utf8",
    timeout: 15_000,
  }).trim();
  const lower = output.toLowerCase();
  assert(
    lower.includes("ansible-core") && lower.includes("ansible-lint"),
    `Expected ansible-core and ansible-lint in output`,
  );
});

Example 2: Create empty playbook

Selenium (Python, 20 lines)

def test_create_empty_playbook(browser_setup, screenshot_on_fail, close_editors):
    driver, _ = browser_setup
    ensure_vscode_ready(driver)
    vscode_run_command(driver, ">ansible.create-empty-playbook")
    wait_displayed(
        driver,
        "//div[contains(@class, 'tab') and contains(., 'Untitled')]",
        timeout=2,
    )
    time.sleep(1)
    view_lines = driver.find_elements("xpath", "//div[@class='view-line']")
    file_lines = [line.text for line in view_lines[:10] if line.text]
    assert len(file_lines) >= 3
    assert any("playbook" in line.lower() for line in file_lines)

WDIO (TypeScript, 15 lines)

Uses the VS Code API directly — no XPath, no DOM scraping:

it("should create an empty playbook via command", async () => {
  const workbench = await browser.getWorkbench();
  await workbench.executeCommand("ansible.create-empty-playbook");
  await browser.pause(3000);

  const text: string = await browser.executeWorkbench((vscode) => {
    return vscode.window.activeTextEditor?.document.getText() ?? "";
  });

  assert.ok(text.length > 0, "Playbook content should not be empty");
  assert.ok(text.split(/\r?\n/).length >= 3);
  await closeActiveEditor(workbench);
});

Example 3: Webview testing (devfile)

Selenium (Python, 25 lines)

Manually traverses nested iframes to find form elements:

def test_devfile_webview(browser_setup, screenshot_on_fail, close_editors):
    driver, _ = browser_setup
    ensure_vscode_ready(driver)
    vscode_run_command(driver, ">Ansible: Create a Devfile")
    find_element_across_iframes(driver, "//form[@id='devfile-form']", retries=10)
    vscode_textfield_interact(driver, "path-url", "~")
    vscode_textfield_interact(driver, "devfile-name", "test")
    vscode_button_click(driver, "create-button")
    overwrite_checkbox = find_element_across_iframes(
        driver, "//vscode-checkbox[@id='overwrite-checkbox']", retries=10,
    )
    overwrite_checkbox.click()
    vscode_button_click(driver, "create-button")
    # ... more interactions ...

Hidden cost: find_element_across_iframes is a 50-line helper that iterates all iframes, nested iframes, catching 4 different exception types, with configurable retries.

WDIO (TypeScript, 15 lines)

WDIO handles iframe context switching automatically:

it("should open the devfile creation webview", async function () {
  this.timeout(60_000);
  await runCommand("ansible.content-creator.create-devfile");
  await browser.pause(3000);

  const workbench = await browser.getWorkbench();
  const webview = await workbench.getWebviewByTitle(/Create Devfile/);
  await webview.open();
  try {
    const form = await $("#devfile-form");
    await form.waitForExist({ timeout: 20_000 });
    assert.ok(await form.isExisting());
  } finally {
    await webview.close();
  }
});

The hidden iceberg: helper code

The Selenium tests look short, but they rely on 1,254 lines of helper code in test/ui/utils/ui_utils.py — functions like ensure_vscode_ready, vscode_run_command, find_element_across_iframes, vscode_textfield_interact, etc.

The WDIO tests have zero helper files for VS Code interaction. The wdio-vscode-service page objects replace all of it:

Selenium helper (custom) WDIO equivalent (built-in)
ensure_vscode_ready() (30 lines) Extension activation is automatic
vscode_run_command() (50 lines) workbench.executeCommand(id)
find_element_across_iframes() (50 lines) workbench.getWebviewByTitle().open()
vscode_textfield_interact() (20 lines) $('#id').setValue(text)
vscode_button_click() (15 lines) $('#id').click()
wait_displayed() (15 lines) element.waitForDisplayed()
get_vscode_file_text() (20 lines) browser.executeWorkbench(vscode => vscode.window.activeTextEditor.document.getText())

Dependencies Added

"@wdio/cli": "^8.46.0",
"@wdio/globals": "^8.46.0",
"@wdio/local-runner": "^8.46.0",
"@wdio/mocha-framework": "^8.46.0",
"@wdio/spec-reporter": "^8.43.0",
"wdio-vscode-service": "^6.1.4",
"express": "^5.2.1",
"ts-node": "^10.9.2"

All are devDependencies. express backs the mock Lightspeed server. ts-node is required by WDIO for TypeScript config and specs.


Summary

Metric Selenium WDIO
Language Python TypeScript (same as extension)
Test target code-server (not VS Code) Real VS Code Electron
Infrastructure Podman container + Selenium Grid None (native Electron)
CI failure rate ~71% TBD (first PR)
Active tests 0 (29 skipped) 30
Lightspeed tests 0 (all skipped) 6 (mock-backed)
Error/timeout tests 0 2
Lines of test helpers 700+ (Python XPath utils) ~190 (mock-server only)
Extension source changes None 2 files (~30 lines)
Container image dependency Yes (selenium-adt) No
Runs locally No (requires podman) Yes (task wdio)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment