Parsing HTML Strings

innerHTML is an element's inner markup, outerHTML adds its own tags, and textContent is plain text that never parses. Assigning innerHTML rebuilds every child, so innerHTML += wipes typed input and listeners. DOMParser parses into a separate, inert Document; XMLSerializer serializes nodes back. Two newer methods:

Parsing the same untrusted string four waysHTMLLive
<pre id="out" style="font-size:12px;margin:0"></pre>
<script>
  const html = '<b>Hi</b><a href="#" onclick="steal()">Go</a><script>steal()<\/script>';
  const box = document.createElement('div');
  box.innerHTML = html;
  const lines = [`innerHTML:  ${box.innerHTML}`];
  if ('setHTML' in box) lines.push(`setHTML:    ${(box.setHTML(html), box.innerHTML)}`);
  const doc = new DOMParser().parseFromString(html, 'text/html');
  lines.push(`DOMParser:  ${doc.body.childElementCount} elements in an inert document`,
    `serialized: ${new XMLSerializer().serializeToString(doc.body.firstChild)}`);
  document.querySelector('#out').textContent = lines.join('\n');
</script>
Browser output of Listing 8.23
Browser output of 23

Inserted scripts never run, but the surviving onclick fires on click: a cross-site scripting hole. Treat innerHTML and ...Unsafe methods as injection sinks (Trusted Types, Sandboxing and Trusted Types; sanitizers, Sanitization and XSS Defence).