zuhtml turns real-world HTML into ordinary R values: character vectors, lists and data frames. It parses the way a browser does, so malformed markup is repaired rather than rejected. You give it a string, raw bytes, a file, a URL or a connection.
A small catalogue page, as it might have been saved from a site. Some
markup is sloppy on purpose: unclosed <li>s, unquoted
attributes, a stray end tag, and a product card without a price.
page <- '
<!DOCTYPE html>
<title>Tea shop</title>
<nav><ul><li><a href="/">Home</a><li><a href="sale/">Sale</a></ul></nav>
<p>Free shipping over 30 EUR</span>
<div class=product>
<h2 class=name>Sencha</h2><span class=price>3.50</span>
<a href="sencha.html">details</a>
</div>
<div class=product>
<h2 class=name>Genmaicha</h2>
<a href="genmaicha.html">details</a>
</div>
<table>
<thead><tr><th>Size<th>Grams</thead>
<tr><td>Small<td>0100
<tr><td>Large<td>0250
</table>'
doc <- html_parse(page, base_url = "https://example.org/shop/")
doc
#> <zuhtml_document>
#> root: <html> with <head>, <body>
#> nodes: 59
#> input: 472 bytes (UTF-8)
#> problems: 1html_read() does the same for a file.
base_url is where the page came from; relative links are
resolved against it.
html_elements() finds every element that matches a CSS
selector. html_element() finds the first match below
each input node, and keeps a missing node where there is none.
That is what keeps extracted columns aligned when some records lack a
field:
cards <- html_elements(doc, ".product")
cards
#> <zuhtml_nodeset[2]>
#> [1] <div class="product">
#> [2] <div class="product">
products <- data.frame(
name = html_text_clean(html_element(cards, ".name")),
price = html_text_clean(html_element(cards, ".price")),
url = html_url(html_element(cards, "a"))
)
products
#> name price url
#> 1 Sencha 3.50 https://example.org/shop/sencha.html
#> 2 Genmaicha <NA> https://example.org/shop/genmaicha.htmlGenmaicha has no price, so it gets NA rather than
shifting the column.
html_text_clean() gives text as a reader wants it;
html_text() gives it exactly as parsed.
html_attr() reads attributes, and
html_serialize() writes nodes back as HTML.
Links, lists and tables have their own extractors:
html_links(doc, absolute = TRUE)
#> text href url
#> 1 Home / https://example.org/
#> 2 Sale sale/ https://example.org/shop/sale/
#> 3 details sencha.html https://example.org/shop/sencha.html
#> 4 details genmaicha.html https://example.org/shop/genmaicha.html
lapply(html_elements(doc, "nav ul"), html_list)
#> [[1]]
#> [1] "Home" "Sale"
html_tables(doc)
#> [[1]]
#> Size Grams
#> 1 Small 0100
#> 2 Large 0250Table columns are character: "0100" keeps its leading
zero. Convert types yourself when you know them, for example with
type.convert().
Real pages nearly always have markup errors, which the parser
repairs. html_problems() lists them:
vignette("selectors"): the supported CSS subset.vignette("tables-and-lists"): how tables and lists are
read.vignette("limits-and-encoding"): resource limits,
encodings, and what zuhtml does not do.