HTML from the web is untrusted input. zuhtml bounds the work and
memory any page can cost, per call, with html_limits():
html_limits()
#> <zuhtml_limits>
#> max_input 16,777,216
#> max_memory 536,870,912
#> max_depth 512
#> max_nodes 4,000,000
#> max_errors 100
#> max_table_cells 1,000,000
#> max_selector_length 16,384max_input is checked before parsing.
max_memory bounds the parser’s native memory and the
document it builds; it is the real guard, since some markup costs far
more memory per byte than other markup.max_depth bounds nesting while parsing. Tree
construction takes time quadratic in nesting depth, so a check after
parsing would come too late: 100,000 nested elements would take 16
seconds. With the limit it fails at once.max_nodes, max_errors (parse problems
kept), max_table_cells and max_selector_length
bound the rest.Exceeding a limit is a zuhtml_limit_error, raised after
every native allocation has been released. It says which limit and by
how much:
err <- tryCatch(
html_parse(strrep("<div>", 1e5)),
zuhtml_limit_error = function(e) e
)
conditionMessage(err)
#> [1] "Elements are nested deeper than max_depth = 512."
err$limit
#> [1] "max_depth"Tighter limits suit a service that parses pages from strangers:
A string is already text: it is used as UTF-8. Raw bytes are decoded,
in order of preference, with a byte-order mark, the
encoding you give, the page’s own <meta>
declaration, or UTF-8:
bytes <- as.raw(c(0x3c, 0x70, 0x3e, 0x63, 0x61, 0x66, 0xe9)) # "<p>caf\xe9"
html_text_clean(html_parse(bytes, encoding = "latin1"))
#> [1] "café"Invalid input is an error, never silently replaced:
The declaration is found as a browser finds it, by scanning the first
1024 bytes for <meta charset> or its
http-equiv form. Labels mean what they mean to browsers, so
iso-8859-1 is read as windows-1252, which makes byte 0x93 a
curly quote rather than a control character:
page <- c(charToRaw("<meta charset=iso-8859-1><p>"), as.raw(0x93),
charToRaw("Quoted"), as.raw(0x94))
doc <- html_parse(page)
html_text_clean(doc)
#> [1] "“Quoted”"
html_info(doc)[c("encoding", "encoding_source")]
#> $encoding
#> [1] "windows-1252"
#>
#> $encoding_source
#> [1] "meta"When you fetch a page, pass the charset from the HTTP
Content-Type header as encoding: it takes
precedence over the page’s declaration, as it does in a browser. A
byte-order mark takes precedence over both; one that contradicts
encoding is an error.
Handle errors by class, never by message text:
tryCatch(
html_elements(html_parse("<p>"), "p:hover"),
zuhtml_selector_error = function(e) paste("unsupported at", e$position)
)
#> [1] "unsupported at 2"See ?zuhtml-conditions for the classes and their
fields.
<script> elements, event-handler attributes and
javascript: URLs. Do not treat
html_serialize() output as safe to embed in another
page.html_read() reads a URL with
base R’s url(), without headers, cookies or retries;
html_url() is string arithmetic and fetches nothing.html_text_clean() follows fixed, documented
rules.