Your RSA-2048 keys break in 2030. Find every one of them before attackers do.
📦 npm

GHSA-5mq8-78gm-pjmq

MEDIUMFix: kepano/defuddle@f154cb7

GHSA-5mq8-78gm-pjmq is a medium-severity (CVSS 6.1) Cross-site Scripting (XSS) vulnerability in defuddle. O3 Security confirms whether GHSA-5mq8-78gm-pjmq is actually reachable in your code before you act, and blocks exploitation at runtime until you patch.

defuddle vulnerable to XSS via unescaped string interpolation in _findContentBySchemaText image tag

Also known asCVE-2026-30830
Published
Mar 6, 2026
Updated
Jul 20, 2026
Affected
1 pkg
Patched
1 / 1
Exploits
None indexed

Real-World Exposure

1 pkg affected
📦defuddle

Real-time download stats are indexed for npm and PyPI packages. This vulnerability affects npm packages — download data is not available via public APIs for these ecosystems.

Description

Summary

The _findContentBySchemaText method in src/defuddle.ts interpolates image src and alt attributes directly into an HTML string without escaping:

html += `<img src="${imageSrc}" alt="${imageAlt}">`;

An attacker can use a " in the alt attribute to break out of the attribute context and inject event handlers. This is a separate vulnerability from the sanitization bypass fixed in f154cb7 — the injection happens during string construction, not in the DOM, so _stripUnsafeElements cannot catch it.

Details

When _findContentBySchemaText finds a sibling image outside the matched content element, it reads the image's src and alt attributes via getAttribute() and interpolates them into a template literal. getAttribute('alt') returns the raw attribute value. If the alt contains ", it terminates the alt attribute in the interpolated HTML string, and subsequent content becomes new attributes (including event handlers).

The recently added _stripUnsafeElements() (commit f154cb7) strips on* attributes from DOM elements, but the alt attribute's name is alt (not on*), so it is preserved with its full value. The onload handler is created by the string interpolation, not present in the original DOM.

PoC

Input HTML:

<!DOCTYPE html>
<html>
<head>
<title>PoC</title>
<script type="application/ld+json">
{"@type": "Article", "text": "Long article text repeated many times to exceed the extracted content word count. Long article text repeated many times to exceed the extracted content word count. Long article text repeated many times to exceed the extracted content word count."}
</script>
</head>
<body>
<article><p>Short.</p></article>
<div class="post-container">
  <p>Extra text to inflate parent word count padding padding padding.</p>
  <div class="post-body">
    Long article text repeated many times to exceed the extracted content word count. Long article text repeated many times to exceed the extracted content word count. Long article text repeated many times to exceed the extracted content word count.
  </div>
  <img width="800" height="600" src="https://example.com/photo.jpg" alt='pwned" onload="alert(document.cookie)'>
</div>
</body>
</html>

Output:

<img src="https://example.com/photo.jpg" alt="pwned" onload="alert(document.cookie)">

The onload event handler is injected as a separate HTML attribute.

Impact

XSS in any application that renders defuddle's HTML output (browser extensions, web clippers, reader modes). The attack requires crafted HTML with schema.org structured data that triggers the _findContentBySchemaText fallback, combined with a sibling image whose alt attribute contains a quote character followed by an event handler.

Suggested Fix

Use DOM API instead of string interpolation:

if (imageSrc) {
    const img = this.doc.createElement('img');
    img.setAttribute('src', imageSrc);
    img.setAttribute('alt', imageAlt);
    html += img.outerHTML;
}

This ensures attribute values are properly escaped by the DOM serializer.

Affected Packages

1 total 1 fixed
EcosystemPackageVulnerable rangeFix
📦npmdefuddleall versions0.9.0

Detection & mitigation playbook

Open-source dependency
  1. Detect

    Scan your dependency tree (package-lock.json, pnpm-lock.yaml, requirements.txt, go.sum, etc.) for defuddle. O3's reachability analysis confirms whether the vulnerable code path is actually invoked in your application, so you act on real exposure instead of every transitive match.

  2. Fix

    Update defuddle to 0.9.0 or later, then make sure no transitive (indirect) dependency still pins the vulnerable range — O3 confirms GHSA-5mq8-78gm-pjmq is resolved across your whole dependency graph.

  3. Workarounds

    If you can't upgrade right away: gate or disable the affected feature, validate untrusted input at the boundary, and avoid passing attacker-controlled data into the vulnerable path. O3's runtime protection blocks exploitation in production as an interim safeguard until the upgrade lands.

  4. How O3 protects you

    O3 pinpoints whether GHSA-5mq8-78gm-pjmq is reachable in your code and exactly where to fix it, then blocks exploitation in production at runtime until the patched version is deployed.

Tailored to GHSA-5mq8-78gm-pjmq. Runtime protection reduces exposure until a permanent patch is applied and verified — it complements patching, it doesn't replace it.

Frequently Asked Questions

### Summary The `_findContentBySchemaText` method in `src/defuddle.ts` interpolates image `src` and `alt` attributes directly into an HTML string without escaping: ```typescript html += `<img src="${imageSrc}" alt="${imageAlt}">`; ``` An attacker can use a `"` in the `alt` attribute to break out of the attribute context and inject event handlers. This is a separate vulnerability from the sanitization bypass fixed in f154cb7 — the injection happens during string construction, not in the DOM, so `_stripUnsafeElements` cannot catch it. ### Details When `_findContentBySchemaText` finds a sibl
O3 Security · Impact-Aware SCA

Is GHSA-5mq8-78gm-pjmq in your dependencies?

O3 detects GHSA-5mq8-78gm-pjmq across npm dependencies and uses function-level reachability to confirm whether the vulnerable code path is actually reachable — not just present. No false positives.