Limits, encodings and safety

library(zuhtml)

Every call runs under limits

HTML from the web is untrusted input. zuhtml bounds the work and memory any page can cost, per call, with html_limits():

html_limits()
#> <zuhtml_limits>
#>   max_input           16,777,216
#>   max_memory          536,870,912
#>   max_depth           512
#>   max_nodes           4,000,000
#>   max_errors          100
#>   max_table_cells     1,000,000
#>   max_selector_length 16,384

Exceeding a limit is a zuhtml_limit_error, raised after every native allocation has been released. It says which limit and by how much:

err <- tryCatch(
  html_parse(strrep("<div>", 1e5)),
  zuhtml_limit_error = function(e) e
)
conditionMessage(err)
#> [1] "Elements are nested deeper than max_depth = 512."
err$limit
#> [1] "max_depth"

Tighter limits suit a service that parses pages from strangers:

strict <- html_limits(max_input = 2 * 1024^2, max_memory = 64 * 1024^2,
                      max_depth = 128)
doc <- html_parse("<p>Small page</p>", limits = strict)

Encodings

A string is already text: it is used as UTF-8. Raw bytes are decoded, in order of preference, with a byte-order mark, the encoding you give, the page’s own <meta> declaration, or UTF-8:

bytes <- as.raw(c(0x3c, 0x70, 0x3e, 0x63, 0x61, 0x66, 0xe9))  # "<p>caf\xe9"
html_text_clean(html_parse(bytes, encoding = "latin1"))
#> [1] "café"

Invalid input is an error, never silently replaced:

try(html_parse(bytes))
#> Error in html_parse(bytes) : The input is not valid UTF-8.

The declaration is found as a browser finds it, by scanning the first 1024 bytes for <meta charset> or its http-equiv form. Labels mean what they mean to browsers, so iso-8859-1 is read as windows-1252, which makes byte 0x93 a curly quote rather than a control character:

page <- c(charToRaw("<meta charset=iso-8859-1><p>"), as.raw(0x93),
          charToRaw("Quoted"), as.raw(0x94))
doc <- html_parse(page)
html_text_clean(doc)
#> [1] "“Quoted”"
html_info(doc)[c("encoding", "encoding_source")]
#> $encoding
#> [1] "windows-1252"
#> 
#> $encoding_source
#> [1] "meta"

When you fetch a page, pass the charset from the HTTP Content-Type header as encoding: it takes precedence over the page’s declaration, as it does in a browser. A byte-order mark takes precedence over both; one that contradicts encoding is an error.

Errors are classed

Handle errors by class, never by message text:

tryCatch(
  html_elements(html_parse("<p>"), "p:hover"),
  zuhtml_selector_error = function(e) paste("unsupported at", e$position)
)
#> [1] "unsupported at 2"

See ?zuhtml-conditions for the classes and their fields.

What zuhtml does not do