String functions join (concat(), ||, string-join()), slice (substring(), substring-before(), substring-after()), measure (string-length()), test (contains(), starts-with(), ends-with()), change case (upper-case(), lower-case()), map characters (translate()) and clean (normalize-space()). Three take regular expressions in XML Schema's dialect: matches(), replace() and tokenize(), with analyze-string() returning the matched and unmatched parts as XML. All are case-sensitive and count characters, not bytes. On BookNest's catalog they do the cleaning a load step needs:
xp 'string-length(//book[1]/summary)' 'string-length(normalize-space(//book[1]/summary))' \
'substring(//book[2]/title, 1, 8)' 'string-join(//book[position() lt 3]/title, "|")' \
'ends-with("ABCDE", "cde")' 'replace((//identifier)[6], "^BN-0*", "")' \
'matches(//sent, "^\d{4}-\d{2}-\d{2}T")' 'tokenize((//sort-name)[2], ",\s*")'string-length(//book[1]/summary) => 81
string-length(normalize-space(//book[1]/summary)) => 75
substring(//book[2]/title, 1, 8) => "Patterns"
string-join(//book[position() lt 3]/title, "|") => "The Quiet Harbor|Patterns of the Deep Web"
ends-with("ABCDE", "cde") => false()
replace((//identifier)[6], "^BN-0*", "") => "6"
matches(//sent, "^\d{4}-\d{2}-\d{2}T") => true()
tokenize((//sort-name)[2], ",\s*") => "Reyes" "Tomas"The first summary is 81 characters as written and 75 after normalize-space() collapses its line break and indentation, which is why every text field should pass through it before it reaches a table. substring() counts from 1, not 0. ends-with("ABCDE", "cde") is false, because case matters (older tables claim otherwise); compare case-insensitively with lower-case() on both sides. XPath's regular expressions have no \b or lookahead, but they do have \i and \c for XML name characters.
String Functions
String functions build, inspect, compare, and transform character strings, including regular-expression matching, splitting, and replacement.| codepoints-to-string((66, 65, 67, 72)) | BACH |
| string-to-codepoints("BACH") | (66, 65, 67, 72) |
| compare("ABC","AAB") | 1 |
| codepoint-equal("ABC","ABC ") | false |
| concat("AB","CD") | ABCD |
| "AB"||"CD" | ABCD |
| string-join(('A','B','C'),'-') | A-B-C |
| substring("ABCDE",2,3) | BCD |
| string-length("ABCDE") | 5 |
| normalize-space('A B C') | A B C |
| normalize-unicode('ABC') | ABC |
| upper-case('Abc') | ABC |
| lower-case('Abc') | abc |
| translate("bar","abc","ABC") | BAr |
| contains("ABCDE","CD") | true |
| starts-with("ABCDE","ABC") | true |
| ends-with("ABCDE","cde") | true |
| substring-before("ABCDE","CDE") | AB |
| substring-after("ABCDE","B") | CDE |
| matches("ABCDE","^A.C.E$") | true |
| replace("abracadabra", "bra", "*") | a*cada* |
| tokenize("AB,CD,E",",") | ('AB','CD','E') |
analyze-string("The cat sat on the mat.", "\w+") matches every run of word characters against the non-matching text around it, and returns:
|
<analyze-string-result xmlns="http://www.w3.org/2005/xpath-functions"> <match>The</match> <non-match> </non-match> <match>cat</match> <non-match> </non-match> <match>sat</match> <non-match> </non-match> <match>on</match> <non-match> </non-match> <match>the</match> <non-match> </non-match> <match>mat</match> <non-match>.</non-match> </analyze-string-result> |
analyze-string("The cat sat on the mat.", "\w+")
<analyze-string-result xmlns="http://www.w3.org/2005/xpath-functions"> <match>The</match> <non-match> </non-match> <match>cat</match> <non-match> </non-match> <match>sat</match> <non-match> </non-match> <match>on</match> <non-match> </non-match> <match>the</match> <non-match> </non-match> <match>mat</match> <non-match>.</non-match> </analyze-string-result>