Selenium → WebDriverIO migration analysis for vscode-ansible
The Selenium UI test suite ran Python pytest against a containerized code-server instance via Selenium Grid. This setup had a 71% CI failure rate, primarily due to infrastructure fragility (container health checks, code-server startup timing, multi-service container coordination). It also tested code-server, not VS Code -- so passing tests did not guarantee real-world correctness.
We replaced it with WebDriverIO (WDIO) using wdio-vscode-service,
which tests against real VS Code (Electron), eliminates all container
infrastructure, and uses the same language (TypeScript) as the extension.
At the time of replacement, all 29 Selenium tests on main were
skipped — zero tests actually executed in CI. WDIO replaces them
with 30 passing tests across 7 spec files, including mock-based
Lightspeed coverage that never existed in CI.
PR: ansible/vscode-ansible#2746
The Selenium container (docker-compose.yml) was introduced on
2026-02-10 in PR #2560. The following table shows weekly CI pass/fail
rates on the main branch from that point forward.
Data source: gh api repos/ansible/vscode-ansible/actions/workflows/ci/runs
(300 completed runs sampled, Feb 10 -- Apr 16, 2026).
| Week | Runs | Pass | Fail | Cancel | Fail Rate | Primary Failure |
|---|---|---|---|---|---|---|
| W09 (Feb 24) | 19 | 18 | 1 | 0 | 5% | builder-image (container build) |
| W10 (Mar 3) | 47 | 43 | 3 | 1 | 6% | builder-image |
| W11 (Mar 10) | 47 | 37 | 4 | 6 | 9% | mixed |
| W12 (Mar 17) | 51 | 43 | 8 | 0 | 16% | builder-image + task ui (selenium) |
| W13 (Mar 24) | 40 | 25 | 15 | 0 | 38% | task ui (selenium) + WSL runner |
| W14 (Mar 31) | 42 | 29 | 13 | 0 | 31% | task ui (selenium) |
| W15 (Apr 7) | 33 | 0 | 33 | 0 | 100% | task ui (selenium) -- every run |
| W16 (Apr 14) | 21 | 0 | 21 | 0 | 100% | task ui (selenium) -- every run |
Key observations:
- Selenium tests were stable for ~3 weeks (W09-W11), then degraded steadily as the container image DOM diverged from the XPath selectors.
- From W15 onward, every single CI run on main failed due to
Selenium, making the
mainbranch permanently red. - The Selenium failure rate over the full period: 98 failures out of 300 runs (33%), with the rate accelerating to 100% in the final two weeks.
- Non-Selenium failures (e2e timeout, WSL runner, container build) accounted for a minority of total failures.
All 29 Selenium tests on main were @pytest.mark.skip — the CI
step still ran (built the container, invoked pytest) but executed
zero actual tests:
| File | Tests | Status | Skip Reason |
|---|---|---|---|
test_00_commands.py |
2 | All skipped | Flaky on CI - extension activation timing |
test_01_dev_webviews.py |
2 | All skipped | Flaky on CI - extension activation timing |
test_02_welcome_webviews.py |
3 | All skipped | Flaky on CI - extension activation timing |
test_03_llm_provider_webview.py |
4 | All skipped | Flaky on CI - extension activation timing |
test_50_mcp_server.py |
4 | All skipped | Flaky on CI - extension activation timing |
test_80_lightspeed.py |
5 | All skipped | AAP-67210 |
test_81_login.py |
8 | All skipped | AAP-67210 |
test_82_lightspeed_trial.py |
1 | All skipped | AAP-67210 |
| Total | 29 | 0 running |
- 15 tests skipped for "Flaky on CI - extension activation timing issues" — the core UI tests (commands, webviews, MCP).
- 14 tests skipped for AAP-67210 — all Lightspeed and SSO tests.
AAP-67210 tracks 14 Lightspeed/SSO tests that were disabled in Selenium. WDIO restores coverage for 5 using a mock server, and adds 2 new tests:
| AAP-67210 Selenium Test | WDIO Status | WDIO Test |
|---|---|---|
test_vscode_widget |
Covered | should have Lightspeed commands registered |
test_vscode_playbook_explanation |
Covered | should open playbook explanation webview |
test_vscode_playbook_generation |
Covered | should open playbook generation webview |
test_vscode_role_generation |
Covered | should open role generation webview |
test_vscode_lightspeed_explorer |
Partial | Explorer view covered via sidebar tests |
test_vscode_trial_button |
Deferred | Excluded from initial PR |
test_unsubed_login |
N/A | SSO flow — tests RH infra, not extension |
test_unsubed_admin_login |
N/A | SSO flow |
test_no_wca_user_login |
N/A | SSO flow |
test_no_wca_admin_login |
N/A | SSO flow |
test_login_page |
N/A | SSO flow |
test_sso_auth_flow |
N/A | SSO flow |
test_admin_portal_error |
N/A | SSO flow |
test_vscode_rhsso_auth_flow |
N/A | SSO flow |
The 8 SSO login tests exercise Red Hat SSO infrastructure (external OAuth pages, live credentials, running Lightspeed backend), not extension code. They were never functional in CI regardless of test framework.
Net-new tests not in AAP-67210: should handle API errors gracefully
and should handle server timeout gracefully — error/timeout resilience
that was never tested.
Host (GitHub Actions runner)
├── Python pytest
│ └── Selenium RemoteWebDriver → http://localhost:4444/wd/hub
│
└── Podman container (ghcr.io/ansible/selenium-adt:main)
├── Selenium Grid (:4444)
├── Firefox/Chrome browser
├── code-server (:8080) ← NOT VS Code
│ └── Ansible extension (from mounted VSIX)
└── VNC (:5999)
| Issue | Impact |
|---|---|
| Health check only validated Selenium Grid, not code-server | 71% CI failure rate |
| code-server != VS Code (DOM differences, API gaps) | Tests may pass but not reflect real user experience |
| Three services in one container | Complex startup coordination |
| Python tests, TypeScript extension | No shared types or utilities |
| Raw XPath selectors + manual iframe traversal | Brittle, breaks on UI changes |
External container image (selenium-adt) |
Build/publish pipeline outside this repo |
| ~10-20 min runtime on CI | Blocks release pipeline |
Host (GitHub Actions runner or local workstation)
├── WDIO test runner (TypeScript)
│ └── wdio-vscode-service
│ ├── Downloads matching VS Code + Chromedriver
│ ├── Launches VS Code Electron with extension
│ └── Provides page objects (Workbench, EditorView, etc.)
│
└── VS Code Electron (real desktop app)
└── Ansible extension (extensionDevelopmentPath)
No containers. No Selenium Grid. No code-server. No Python.
| Item | Description |
|---|---|
test/ui/ (entire directory) |
Python pytest tests, fixtures, XPath helpers (700+ lines) |
docker-compose.yml |
Selenium/code-server container orchestration |
podman-compose (pyproject.toml) |
Python dependency for container management |
selenium (pyproject.toml) |
Python dependency for WebDriver |
| Container health checks | Startup coordination for 3 in-container services |
| VNC debugging | Replaced by native VS Code and screenshots |
| File | Purpose | Lines |
|---|---|---|
wdio.conf.ts |
WDIO configuration (VS Code service, Mocha, specs) | 57 |
test/wdio/smoke.spec.ts |
Extension activation, activity bar, commands, language detection | 99 |
test/wdio/terminal.spec.ts |
adt --version availability, terminal API |
43 |
test/wdio/commands.spec.ts |
ansible.create-empty-playbook command |
71 |
test/wdio/webviews.spec.ts |
Devfile, devcontainer, welcome page, LLM settings | 219 |
test/wdio/mcp.spec.ts |
MCP enable/disable via commands and settings | 174 |
test/wdio/sidebar.spec.ts |
ADT sidebar, welcome page link, settings propagation | 90 |
test/wdio/lightspeed.spec.ts |
Mock-based Lightspeed tests (new coverage) | 303 |
test/wdio/mock-server.ts |
Express mock server for Lightspeed API | 190 |
test/wdio/fixtures/ |
Playbook files and VS Code settings for tests | — |
scripts/install-test-extensions.mjs |
Downloads VS Code + installs dependency extensions | 43 |
test/wdio/tsconfig.json |
TypeScript config for WDIO specs | — |
| File | Change |
|---|---|
src/extension.ts |
Added ansible.lightspeed.mockSession command for test session injection |
src/features/lightspeed/lightSpeedOAuthProvider.ts |
Added setMockSession() method to write fake auth sessions |
The mockSession command is not contributed in package.json, has no
UI surface, and only writes to the extension's local secret store. It
exists solely to let WDIO tests bypass SSO without live credentials.
| # | Selenium test | WDIO spec | Approach |
|---|---|---|---|
| 1 | test_terminal |
terminal.spec.ts |
execSync("adt --version") + terminal API |
| 2 | test_create_empty_playbook |
commands.spec.ts |
executeWorkbench() + editor content |
| 3 | test_devfile_webview |
webviews.spec.ts |
Webview context switch + content polling |
| 4 | test_devcontainer_webview |
webviews.spec.ts |
Same pattern |
| 5 | test_header_and_subtitle |
webviews.spec.ts |
Welcome page header assertion |
| 6 | test_header_subtitle |
webviews.spec.ts |
Welcome page subtitle assertion |
| 7 | test_mcp_section |
webviews.spec.ts |
MCP section in welcome page |
| 8-11 | test_llm_provider_* (x4) |
webviews.spec.ts |
LLM settings webview (list, edit, connect) |
| 12-15 | test_mcp_server_* (x4) |
mcp.spec.ts |
Enable/disable via commands and settings |
| 16 | test_sidebar_nav |
sidebar.spec.ts |
Activity bar + sidebar view |
| # | Test | Spec file | What it validates | Why Selenium couldn't |
|---|---|---|---|---|
| 17 | VS Code launch | smoke.spec.ts |
Session starts successfully | No validation of this |
| 18 | Activity bar icon | smoke.spec.ts |
Ansible icon visible in activity bar | XPath-dependent |
| 19 | Extension activation | smoke.spec.ts |
isActive === true via VS Code API |
No vscode API access |
| 20 | Command registration | smoke.spec.ts |
ansible.* commands exist |
Couldn't enumerate |
| 21 | Language detection | smoke.spec.ts |
.ansible.yml → languageId "ansible" |
No API; indirect |
| 22 | Welcome page link | sidebar.spec.ts |
Sidebar opens welcome page | Never tested |
| 23 | Settings propagation | sidebar.spec.ts |
Runtime config change round-trip | Pre-baked only |
| 24 | Terminal API | terminal.spec.ts |
Create and dispose terminal | Only scraped DOM |
| 25 | Lightspeed commands | lightspeed.spec.ts |
Commands registered after mock login | Required live SSO |
| 26 | Playbook explanation | lightspeed.spec.ts |
Mock API → explanation webview | Required live SSO |
| 27 | Playbook generation | lightspeed.spec.ts |
Mock API → generation webview | Required live SSO |
| 28 | Role generation | lightspeed.spec.ts |
Mock API → role webview | Required live SSO |
| 29 | API error handling | lightspeed.spec.ts |
500 → graceful degradation | Never tested |
| 30 | Timeout handling | lightspeed.spec.ts |
Slow server → timeout behavior | Never tested |
Total: 30 tests vs 16 active Selenium tests (88% increase).
The Lightspeed tests (#25-30) are the most significant gain: they test end-to-end UI flows that were permanently skipped in Selenium because they required live Red Hat SSO credentials and a running Lightspeed backend. WDIO tests use a local Express mock server and injected authentication sessions instead.
A single Express server (test/wdio/mock-server.ts) stubs all
Lightspeed API endpoints with configurable canned responses:
POST /api/v0/ai/completions/ → inline suggestions
POST /api/v0/ai/explanations/ → playbook explanations
POST /api/v0/ai/generations/ → playbook/role generation
GET /api/v0/me/ → user info
Individual tests can override any endpoint via setResponse(endpoint, status, body). Overrides are cleared between tests with
resetResponses().
Instead of real SSO, tests call a ansible.lightspeed.mockSession
command that injects a synthetic auth session into the extension's
secret store:
await browser.executeWorkbench(async (vscode) => {
await vscode.commands.executeCommand(
"ansible.lightspeed.mockSession",
{ accessToken: "mock", accountId: "test", accountLabel: "Test" },
);
});WDIO's executeWorkbench runs code inside the VS Code process with
full vscode API access. This replaces XPath/CSS selectors for
assertions that aren't about the DOM:
- Check extension activation state
- Read/write configuration values
- Execute commands by ID
- Open files and check language IDs
- Query editor content
WDIO's switchToFrame() handles VS Code webview iframe traversal
automatically. A polling helper waitForWebviewText() retries until
Vue hydration completes, avoiding the race conditions that plagued the
Selenium iframe navigation.
- name: Configure podman
run: ...
- name: Run UI tests (Selenium)
run: task ui # podman-compose up → pytest → podman-compose down- name: Run UI tests (WDIO)
run: task wdio # pretest:wdio → xvfb-run npx wdioWDIO runs on the same Linux runner as unit and e2e tests. No containers,
no podman, no extra runner. Uses xvfb-run for headless display (same
as the existing e2e tests).
task wdio # full suite
npx wdio --spec smoke # single spec
npx wdio --spec lightspeed # lightspeed tests onlyRequires xvfb-run on Linux (Wayland users: sudo dnf install xorg-x11-server-Xvfb).
The test/e2e/ suite (@vscode/test-cli + Mocha) tests the extension
API layer inside VS Code Electron:
| E2E Coverage | UI Coverage (WDIO) | Overlap |
|---|---|---|
| Hover providers (with/without EE) | — | None |
| ansible-lint diagnostics | — | None |
| YAML validation diagnostics | — | None |
| Outside-workspace LS | — | None |
| MCP activation (API-level) | MCP enable/disable (UI-level) | Complementary |
| — | Terminal, commands | None |
| — | Webviews (devfile, welcome, LLM) | None |
| — | Lightspeed flows | None |
The two suites test different layers with minimal overlap. Both are kept. WDIO replaced only the Selenium UI suite.
| Item | Why |
|---|---|
| Full SSO login flow (external RH SSO page) | Tests Red Hat infrastructure, not the extension |
| Admin portal error test | Same |
| Trial button test | Deferred from initial PR |
| Inline suggestion ghost text | Technically possible but tricky; deferred |
The SSO/OAuth tests were skipped in Selenium (AAP-67210) and remain out of scope. They test Red Hat infrastructure, not extension behavior.
Selenium (Python, 35 lines)
The test opens the command palette via XPath, sends a terminal command
through sendSequence, then scrapes the terminal DOM to find output:
def test_terminal(browser_setup, screenshot_on_fail, close_editors):
driver, _ = browser_setup
ensure_vscode_ready(driver) # 120s wait
vscode_run_command(driver, ">workbench.action.terminal.new") # F1 → XPath
vscode_run_command(driver, ">workbench.action.terminal.focus")
time.sleep(3) # arbitrary sleep
vscode_run_command(
driver, ">workbench.action.terminal.sendSequence",
"adt --version\\n"
)
output = None
def check_output():
nonlocal output
output = driver.find_element(
by="xpath", value="//div[@class='terminal-xterm-host']"
)
return "tox-ansible" in output.text
errors = [NoSuchElementException, ElementNotInteractableException]
wait = WebDriverWait(driver, timeout=10, poll_frequency=0.5,
ignored_exceptions=errors)
wait.until(lambda _: check_output())
text = output.text
missing = [pkg for pkg in EXPECTED_ADT_PACKAGES if pkg not in text]
assert not missingHidden cost: vscode_run_command is a 50-line helper that clicks the
command center via XPath, retries up to 6 times with time.sleep(1),
and manually sends keystrokes to the input box.
WDIO (TypeScript, 10 lines)
Runs adt --version directly in the test process — no DOM scraping:
it("should have adt available", async function () {
this.timeout(30_000);
const output = execSync("adt --version", {
encoding: "utf8",
timeout: 15_000,
}).trim();
const lower = output.toLowerCase();
assert(
lower.includes("ansible-core") && lower.includes("ansible-lint"),
`Expected ansible-core and ansible-lint in output`,
);
});Selenium (Python, 20 lines)
def test_create_empty_playbook(browser_setup, screenshot_on_fail, close_editors):
driver, _ = browser_setup
ensure_vscode_ready(driver)
vscode_run_command(driver, ">ansible.create-empty-playbook")
wait_displayed(
driver,
"//div[contains(@class, 'tab') and contains(., 'Untitled')]",
timeout=2,
)
time.sleep(1)
view_lines = driver.find_elements("xpath", "//div[@class='view-line']")
file_lines = [line.text for line in view_lines[:10] if line.text]
assert len(file_lines) >= 3
assert any("playbook" in line.lower() for line in file_lines)WDIO (TypeScript, 15 lines)
Uses the VS Code API directly — no XPath, no DOM scraping:
it("should create an empty playbook via command", async () => {
const workbench = await browser.getWorkbench();
await workbench.executeCommand("ansible.create-empty-playbook");
await browser.pause(3000);
const text: string = await browser.executeWorkbench((vscode) => {
return vscode.window.activeTextEditor?.document.getText() ?? "";
});
assert.ok(text.length > 0, "Playbook content should not be empty");
assert.ok(text.split(/\r?\n/).length >= 3);
await closeActiveEditor(workbench);
});Selenium (Python, 25 lines)
Manually traverses nested iframes to find form elements:
def test_devfile_webview(browser_setup, screenshot_on_fail, close_editors):
driver, _ = browser_setup
ensure_vscode_ready(driver)
vscode_run_command(driver, ">Ansible: Create a Devfile")
find_element_across_iframes(driver, "//form[@id='devfile-form']", retries=10)
vscode_textfield_interact(driver, "path-url", "~")
vscode_textfield_interact(driver, "devfile-name", "test")
vscode_button_click(driver, "create-button")
overwrite_checkbox = find_element_across_iframes(
driver, "//vscode-checkbox[@id='overwrite-checkbox']", retries=10,
)
overwrite_checkbox.click()
vscode_button_click(driver, "create-button")
# ... more interactions ...Hidden cost: find_element_across_iframes is a 50-line helper that
iterates all iframes, nested iframes, catching 4 different exception
types, with configurable retries.
WDIO (TypeScript, 15 lines)
WDIO handles iframe context switching automatically:
it("should open the devfile creation webview", async function () {
this.timeout(60_000);
await runCommand("ansible.content-creator.create-devfile");
await browser.pause(3000);
const workbench = await browser.getWorkbench();
const webview = await workbench.getWebviewByTitle(/Create Devfile/);
await webview.open();
try {
const form = await $("#devfile-form");
await form.waitForExist({ timeout: 20_000 });
assert.ok(await form.isExisting());
} finally {
await webview.close();
}
});The hidden iceberg: helper code
The Selenium tests look short, but they rely on 1,254 lines of
helper code in test/ui/utils/ui_utils.py — functions like
ensure_vscode_ready, vscode_run_command,
find_element_across_iframes, vscode_textfield_interact, etc.
The WDIO tests have zero helper files for VS Code interaction.
The wdio-vscode-service page objects replace all of it:
| Selenium helper (custom) | WDIO equivalent (built-in) |
|---|---|
ensure_vscode_ready() (30 lines) |
Extension activation is automatic |
vscode_run_command() (50 lines) |
workbench.executeCommand(id) |
find_element_across_iframes() (50 lines) |
workbench.getWebviewByTitle().open() |
vscode_textfield_interact() (20 lines) |
$('#id').setValue(text) |
vscode_button_click() (15 lines) |
$('#id').click() |
wait_displayed() (15 lines) |
element.waitForDisplayed() |
get_vscode_file_text() (20 lines) |
browser.executeWorkbench(vscode => vscode.window.activeTextEditor.document.getText()) |
"@wdio/cli": "^8.46.0",
"@wdio/globals": "^8.46.0",
"@wdio/local-runner": "^8.46.0",
"@wdio/mocha-framework": "^8.46.0",
"@wdio/spec-reporter": "^8.43.0",
"wdio-vscode-service": "^6.1.4",
"express": "^5.2.1",
"ts-node": "^10.9.2"All are devDependencies. express backs the mock Lightspeed server.
ts-node is required by WDIO for TypeScript config and specs.
| Metric | Selenium | WDIO |
|---|---|---|
| Language | Python | TypeScript (same as extension) |
| Test target | code-server (not VS Code) | Real VS Code Electron |
| Infrastructure | Podman container + Selenium Grid | None (native Electron) |
| CI failure rate | ~71% | TBD (first PR) |
| Active tests | 0 (29 skipped) | 30 |
| Lightspeed tests | 0 (all skipped) | 6 (mock-backed) |
| Error/timeout tests | 0 | 2 |
| Lines of test helpers | 700+ (Python XPath utils) | ~190 (mock-server only) |
| Extension source changes | None | 2 files (~30 lines) |
| Container image dependency | Yes (selenium-adt) |
No |
| Runs locally | No (requires podman) | Yes (task wdio) |