# Files changed · Pull request \#1 · versecafe/webls

[View on GitCafe](https://git.cafe/versecafe/webls/pull/1/diffs)

Repository: [versecafe/webls](https://git.cafe/versecafe/webls)

Visibility: public

[webls v2](https://git.cafe/versecafe/webls/pull/1.md)

Version: 1

Head commit: 46e9a680baecefc5a43edb883ba7e05f65fbf61b

Base commit: defe0263d40e7e5d835dff01ecf14b31c532fb68

Comparison base: defe0263d40e7e5d835dff01ecf14b31c532fb68

Showing up to 20 entries from this revision. If the pull request changes, reopen the first page before continuing.

## \.github/workflows/test.yml

Status: modified; +24; −2

```diff
@@ -3,13 +3,19 @@
 on:
   push:
     branches:
-      - master
       - main
   pull_request:
 
 jobs:
   test:
     runs-on: ubuntu-latest
+    strategy:
+      matrix:
+        target:
+          - { name: erlang, runtime: "" }
+          - { name: javascript, runtime: node }
+          - { name: javascript, runtime: bun }
+          - { name: javascript, runtime: deno }
     steps:
       - uses: actions/checkout@v6
       - uses: erlef/setup-beam@v1
@@ -17,6 +23,22 @@
           otp-version: "28.3.0"
           gleam-version: "1.14.0"
           rebar3-version: "3.26.0"
+      - uses: oven-sh/setup-bun@v2
+        if: matrix.target.runtime == 'bun'
+      - uses: denoland/setup-deno@v2
+        if: matrix.target.runtime == 'deno'
+        with:
+          deno-version: v2.x
       - run: gleam deps download
+      - run: gleam test --target ${{ matrix.target.name }} ${{ matrix.target.runtime && format('--runtime {0}', matrix.target.runtime) || '' }}
+
+  format:
+    runs-on: ubuntu-latest
+    steps:
+      - uses: actions/checkout@v4
+      - uses: erlef/setup-beam@v1
+        with:
-      - run: gleam test
+          otp-version: "28.3.0"
+          gleam-version: "1.14.0"
+          rebar3-version: "3.26.0"
       - run: gleam format --check src test
```

## README.md

Status: modified; +6; −7

````diff
@@ -78,17 +78,16 @@
 
 ## Utility Support
 
+| Type       | Builder Functions | to_string | from_string |
+| ---------- | ----------------- | --------- | ----------- |
+| Sitemap    | Complete          | Complete  | Complete    |
+| RSS v2.0   | Complete          | Complete  | Complete    |
+| Robots.txt | Complete          | Complete  | Complete    |
+| Atom       | Complete          | Complete  | None        |
-| Type       | to_string | Builder Functions | Validators |
-| ---------- | --------- | ----------------- | ---------- |
-| Sitemap    | Complete  | Complete          | None       |
-| RSS v2.0   | Complete  | Complete          | None       |
-| Robots.txt | Complete  | Complete          | None       |
-| Atom       | Complete  | Complete          | None       |
 
 ## Development
 
 ```sh
-gleam run   # Run the project
 gleam test  # Run the tests
 ```
 
````

## codebook.toml

Status: modified; +1; −1

```diff
@@ -7,7 +7,7 @@
     "webls", "rss", "googlebot", "bingbot",
     # code words
     "ttl", "perma", "rebar", "otp", "erlef", "gleeunit",
-    "simplifile",
+    "simplifile", "deno",
     # other words
     "versecafe"
 ]
```

## gleam.toml

Status: modified; +5; −1

```diff
@@ -1,5 +1,5 @@
 name = "webls"
-version = "1.6.1"
+version = "2.0.0"
 
 description = "A simple web utility library for RSS feeds, Sitemaps, Robots.txt, etc."
 licences = ["Apache-2.0"]
@@ -8,7 +8,11 @@
 [dependencies]
 gleam_stdlib = ">= 0.34.0 and < 2.0.0"
 gleam_time = ">= 1.6.0 and < 2.0.0"
+parsed_it = ">= 0.1.1 and < 0.2.0"
 
 [dev-dependencies]
 gleeunit = ">= 1.0.0 and < 2.0.0"
 simplifile = ">= 2.3.2 and < 3.0.0"
+
+[javascript.deno]
+allow_read = true
```

## manifest.toml

Status: modified; +3; −1

```diff
@@ -4,8 +4,9 @@
 packages = [
   { name = "filepath", version = "1.1.2", build_tools = ["gleam"], requirements = ["gleam_stdlib"], otp_app = "filepath", source = "hex", outer_checksum = "B06A9AF0BF10E51401D64B98E4B627F1D2E48C154967DA7AF4D0914780A6D40A" },
   { name = "gleam_stdlib", version = "0.68.1", build_tools = ["gleam"], requirements = [], otp_app = "gleam_stdlib", source = "hex", outer_checksum = "F7FAEBD8EF260664E86A46C8DBA23508D1D11BB3BCC6EE1B89B3BC3E5C83FF1E" },
-  { name = "gleam_time", version = "1.6.0", build_tools = ["gleam"], requirements = ["gleam_stdlib"], otp_app = "gleam_time", source = "hex", outer_checksum = "0DF3834D20193F0A38D0EB21F0A78D48F2EC276C285969131B86DF8D4EF9E762" },
+  { name = "gleam_time", version = "1.7.0", build_tools = ["gleam"], requirements = ["gleam_stdlib"], otp_app = "gleam_time", source = "hex", outer_checksum = "56DB0EF9433826D3B99DB0B4AF7A2BFED13D09755EC64B1DAAB46F804A9AD47D" },
   { name = "gleeunit", version = "1.9.0", build_tools = ["gleam"], requirements = ["gleam_stdlib"], otp_app = "gleeunit", source = "hex", outer_checksum = "DA9553CE58B67924B3C631F96FE3370C49EB6D6DC6B384EC4862CC4AAA718F3C" },
+  { name = "parsed_it", version = "0.1.1", build_tools = ["gleam"], requirements = ["gleam_stdlib"], otp_app = "parsed_it", source = "hex", outer_checksum = "9F8BA3C634FEA847AD195E3322FD1DA51980F57C4171B02DCF069C6FC807944A" },
   { name = "simplifile", version = "2.3.2", build_tools = ["gleam"], requirements = ["filepath", "gleam_stdlib"], otp_app = "simplifile", source = "hex", outer_checksum = "E049B4DACD4D206D87843BCF4C775A50AE0F50A52031A2FFB40C9ED07D6EC70A" },
 ]
 
@@ -13,4 +14,5 @@
 gleam_stdlib = { version = ">= 0.34.0 and < 2.0.0" }
 gleam_time = { version = ">= 1.6.0 and < 2.0.0" }
 gleeunit = { version = ">= 1.0.0 and < 2.0.0" }
+parsed_it = { version = ">= 0.1.1 and < 0.2.0" }
 simplifile = { version = ">= 2.3.2 and < 3.0.0" }
```

## src/webls/robots.gleam

Status: modified; +203; −12

````diff
@@ -1,16 +1,57 @@
+//// Functions for building and parsing robots.txt files.
+////
+//// ## Building a robots.txt
+////
+//// ```gleam
+//// import webls/robots
+////
+//// robots.config("https://example.com/sitemap.xml")
+//// |> robots.with_config_robot(
+////   robots.robot("*")
+////   |> robots.with_robot_disallowed_route("/admin/")
+//// )
+//// |> robots.to_string
+//// ```
+////
+//// ## Parsing a robots.txt
+////
+//// ```gleam
+//// import webls/robots
+////
+//// let assert Ok(config) = robots.from_string(robots_txt_content)
+//// // Access config.sitemap_url and config.robots
+//// ```
+////
+//// The parser handles comments, extra whitespace, and case-insensitive
+//// directives. Unknown directives are ignored. Malformed lines (missing `:`)
+//// return an error.
+
 import gleam/list
+import gleam/option.{type Option, None, Some}
 import gleam/result
+import gleam/string
 
 // Stringify ------------------------------------------------------------------
 
+/// Converts a RobotsConfig to a robots.txt formatted string.
+///
+/// The output format follows the standard robots.txt specification:
+/// - Sitemap directive at the top (if present)
+/// - User-agent blocks separated by blank lines
+/// - Allow directives followed by Disallow directives for each agent
 pub fn to_string(config: RobotsConfig) -> String {
+  let sitemap_section = case config.sitemap_url {
+    Some(url) -> "Sitemap: " <> url <> "\n\n"
+    None -> ""
+  }
+
+  let robots_section =
+    config.robots
+    |> list.map(fn(robot) { robot |> robot_to_string })
+    |> list.reduce(fn(acc, line) { acc <> "\n\n" <> line })
+    |> result.unwrap("")
+
-  "Sitemap: "
-  <> config.sitemap_url
-  <> "\n\n"
-  <> config.robots
-  |> list.map(fn(robot) { robot |> robot_to_string })
-  |> list.reduce(fn(acc, line) { acc <> "\n\n" <> line })
-  |> result.unwrap("")
+  sitemap_section <> robots_section
 }
 
 fn robot_to_string(robot: Robot) -> String {
@@ -28,11 +69,24 @@
   |> result.unwrap("")
 }
 
-// Builder Patern -------------------------------------------------------------
+// Builder Pattern ------------------------------------------------------------
 
 /// Creates a robots config with a sitemap url
 pub fn config(sitemap_url: String) -> RobotsConfig {
+  RobotsConfig(sitemap_url: Some(sitemap_url), robots: [])
+}
+
+/// Creates a robots config without a sitemap url
+pub fn config_without_sitemap() -> RobotsConfig {
+  RobotsConfig(sitemap_url: None, robots: [])
+}
+
+/// Sets the sitemap url on a robots config
+pub fn with_config_sitemap(
+  config: RobotsConfig,
-  RobotsConfig(sitemap_url: sitemap_url, robots: [])
+  sitemap_url: String,
+) -> RobotsConfig {
+  RobotsConfig(..config, sitemap_url: Some(sitemap_url))
 }
 
 /// Adds a list of robots to the robots config
@@ -58,7 +112,7 @@
   Robot(..robot, allowed_routes: list.flatten([robot.allowed_routes, routes]))
 }
 
-/// Adds a allowed route to the robot policy
+/// Adds an allowed route to the robot policy
 pub fn with_robot_allowed_route(robot: Robot, route: String) -> Robot {
   Robot(..robot, allowed_routes: [route, ..robot.allowed_routes])
 }
@@ -81,8 +135,8 @@
 /// The configuration for a robots.txt file
 pub type RobotsConfig {
   RobotsConfig(
+    /// The optional url of the sitemap for crawlers to use
-    /// The url of the sitemap for crawlers to use
+    sitemap_url: Option(String),
-    sitemap_url: String,
     /// A list of robot policies
     robots: List(Robot),
   )
@@ -99,3 +153,140 @@
     disallowed_routes: List(String),
   )
 }
+
+/// Error returned when parsing a malformed robots.txt line
+pub type RobotsParseError {
+  /// A line could not be parsed as a valid directive (missing `:`)
+  InvalidDirective(line: String)
+}
+
+// Parse ----------------------------------------------------------------------
+
+/// Parses a robots.txt string into a RobotsConfig.
+///
+/// The parser handles:
+/// - Case-insensitive directives (e.g., `USER-AGENT`, `user-agent`)
+/// - Comments (lines starting with `#` or inline `# comment`)
+/// - Extra whitespace around directives and values
+/// - Unknown directives (silently ignored)
+///
+/// Returns an error if a non-empty, non-comment line is malformed (missing `:`).
+/// An empty config (no sitemap, no robots) is valid.
+/// Directives appearing before any `User-agent:` line are ignored.
+pub fn from_string(input: String) -> Result(RobotsConfig, RobotsParseError) {
+  let lines =
+    input
+    |> string.split("\n")
+    |> list.map(strip_comment)
+    |> list.map(string.trim)
+    |> list.filter(fn(line) { line != "" })
+
+  case validate_lines(lines) {
+    Error(e) -> Error(e)
+    Ok(_) -> {
+      let sitemap_url = find_sitemap(lines)
+      let robot_lines = list.filter(lines, fn(line) { !is_sitemap_line(line) })
+      let robots = parse_robots(robot_lines, [], None)
+      Ok(RobotsConfig(sitemap_url: sitemap_url, robots: robots))
+    }
+  }
+}
+
+/// Validates that all lines are valid directives (contain `:`)
+fn validate_lines(lines: List(String)) -> Result(Nil, RobotsParseError) {
+  case lines {
+    [] -> Ok(Nil)
+    [line, ..rest] ->
+      case string.contains(line, ":") {
+        True -> validate_lines(rest)
+        False -> Error(InvalidDirective(line))
+      }
+  }
+}
+
+/// Strips inline comments from a line (everything after `#`)
+fn strip_comment(line: String) -> String {
+  case string.split_once(line, "#") {
+    Ok(#(before, _)) -> before
+    Error(_) -> line
+  }
+}
+
+/// Splits a directive line into key and value on the first `:`
+fn split_directive(line: String) -> Result(#(String, String), Nil) {
+  case string.split_once(line, ":") {
+    Ok(#(key, value)) -> Ok(#(string.trim(key), string.trim(value)))
+    Error(_) -> Error(Nil)
+  }
+}
+
+fn is_sitemap_line(line: String) -> Bool {
+  case split_directive(line) {
+    Ok(#(key, _)) -> string.lowercase(key) == "sitemap"
+    Error(_) -> False
+  }
+}
+
+fn find_sitemap(lines: List(String)) -> Option(String) {
+  lines
+  |> list.find(is_sitemap_line)
+  |> result.map(fn(line) {
+    case split_directive(line) {
+      Ok(#(_, value)) -> value
+      Error(_) -> ""
+    }
+  })
+  |> option.from_result
+}
+
+fn parse_robots(
+  lines: List(String),
+  acc: List(Robot),
+  current: Option(Robot),
+) -> List(Robot) {
+  case lines {
+    [] ->
+      case current {
+        Some(r) -> list.reverse([r, ..acc])
+        None -> list.reverse(acc)
+      }
+    [line, ..rest] -> {
+      case split_directive(line) {
+        Ok(#(key, value)) -> {
+          let lower_key = string.lowercase(key)
+          case lower_key {
+            "user-agent" -> {
+              let new_robot = Robot(value, [], [])
+              case current {
+                Some(r) -> parse_robots(rest, [r, ..acc], Some(new_robot))
+                None -> parse_robots(rest, acc, Some(new_robot))
+              }
+            }
+            _ ->
+              case current {
+                Some(r) -> {
+                  let updated = parse_directive(lower_key, value, r)
+                  parse_robots(rest, acc, Some(updated))
+                }
+                None -> parse_robots(rest, acc, None)
+              }
+          }
+        }
+        Error(_) -> parse_robots(rest, acc, current)
+      }
+    }
+  }
+}
+
+fn parse_directive(key: String, value: String, robot: Robot) -> Robot {
+  case key {
+    "allow" ->
+      Robot(..robot, allowed_routes: list.append(robot.allowed_routes, [value]))
+    "disallow" ->
+      Robot(
+        ..robot,
+        disallowed_routes: list.append(robot.disallowed_routes, [value]),
+      )
+    _ -> robot
+  }
+}
````

## src/webls/rss.gleam

Status: modified; +319; −0

```diff
@@ -1,9 +1,12 @@
+import gleam/dynamic/decode
 import gleam/int
 import gleam/list
 import gleam/option.{type Option, None, Some}
 import gleam/result
+import gleam/string
 import gleam/time/calendar
 import gleam/time/timestamp.{type Timestamp}
+import parsed_it/xml
 
 // Stringify ------------------------------------------------------------------
 
@@ -613,3 +616,319 @@
   Saturday
   Sunday
 }
+
+// Decoders -------------------------------------------------------------------
+
+/// Parses an RSS XML string into a list of RssChannels
+pub fn from_string(
+  rss_xml: String,
+) -> Result(List(RssChannel), xml.XmlDecodeError) {
+  xml.parse(from: rss_xml, using: rss_decoder())
+}
+
+fn rss_decoder() -> decode.Decoder(List(RssChannel)) {
+  // rss -> channel (single or list)
+  use channels <- decode.field(
+    "channel",
+    decode.one_of(decode.list(channel_decoder()), [
+      channel_decoder() |> decode.map(fn(ch) { [ch] }),
+    ]),
+  )
+  decode.success(channels)
+}
+
+fn channel_decoder() -> decode.Decoder(RssChannel) {
+  use title <- decode.field("title", text_decoder())
+  use link <- decode.field("link", text_decoder())
+  use description <- decode.field("description", text_decoder())
+  use language <- decode.optional_field(
+    "language",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use copyright <- decode.optional_field(
+    "copyright",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use managing_editor <- decode.optional_field(
+    "managingEditor",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use web_master <- decode.optional_field(
+    "webMaster",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use pub_date <- decode.optional_field(
+    "pubDate",
+    None,
+    decode.optional(timestamp_decoder()),
+  )
+  use last_build_date <- decode.optional_field(
+    "lastBuildDate",
+    None,
+    decode.optional(timestamp_decoder()),
+  )
+  use categories <- decode.optional_field("category", [], categories_decoder())
+  use generator <- decode.optional_field(
+    "generator",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use docs <- decode.optional_field(
+    "docs",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use cloud <- decode.optional_field(
+    "cloud",
+    None,
+    decode.optional(cloud_decoder()),
+  )
+  use ttl <- decode.optional_field(
+    "ttl",
+    None,
+    decode.optional(int_text_decoder()),
+  )
+  use image <- decode.optional_field(
+    "image",
+    None,
+    decode.optional(image_decoder()),
+  )
+  use text_input <- decode.optional_field(
+    "textInput",
+    None,
+    decode.optional(text_input_decoder()),
+  )
+  use skip_hours <- decode.optional_field("skipHours", [], skip_hours_decoder())
+  use skip_days <- decode.optional_field("skipDays", [], skip_days_decoder())
+  use items <- decode.optional_field(
+    "item",
+    [],
+    decode.one_of(decode.list(item_decoder()), [
+      item_decoder() |> decode.map(fn(item) { [item] }),
+    ]),
+  )
+  decode.success(RssChannel(
+    title:,
+    link:,
+    description:,
+    language:,
+    copyright:,
+    managing_editor:,
+    web_master:,
+    pub_date:,
+    last_build_date:,
+    categories:,
+    generator:,
+    docs:,
+    cloud:,
+    ttl:,
+    image:,
+    text_input:,
+    skip_hours:,
+    skip_days:,
+    items:,
+  ))
+}
+
+fn item_decoder() -> decode.Decoder(RssItem) {
+  use title <- decode.field("title", text_decoder())
+  use description <- decode.field("description", text_decoder())
+  use link <- decode.optional_field(
+    "link",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use author <- decode.optional_field(
+    "author",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use comments <- decode.optional_field(
+    "comments",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use source <- decode.optional_field(
+    "source",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use pub_date <- decode.optional_field(
+    "pubDate",
+    None,
+    decode.optional(timestamp_decoder()),
+  )
+  use categories <- decode.optional_field("category", [], categories_decoder())
+  use enclosure <- decode.optional_field(
+    "enclosure",
+    None,
+    decode.optional(enclosure_decoder()),
+  )
+  use guid <- decode.optional_field(
+    "guid",
+    None,
+    decode.optional(guid_decoder()),
+  )
+  decode.success(RssItem(
+    title:,
+    description:,
+    link:,
+    author:,
+    comments:,
+    source:,
+    pub_date:,
+    categories:,
+    enclosure:,
+    guid:,
+  ))
+}
+
+fn text_decoder() -> decode.Decoder(String) {
+  decode.at(["$text"], decode.string)
+}
+
+fn int_text_decoder() -> decode.Decoder(Int) {
+  use str <- decode.then(text_decoder())
+  case int.parse(str) {
+    Ok(i) -> decode.success(i)
+    Error(_) -> decode.failure(0, "Int")
+  }
+}
+
+fn timestamp_decoder() -> decode.Decoder(Timestamp) {
+  use date_str <- decode.then(text_decoder())
+  case timestamp.parse_rfc3339(date_str) {
+    Ok(ts) -> decode.success(ts)
+    Error(_) -> decode.failure(timestamp.from_unix_seconds(0), "Timestamp")
+  }
+}
+
+fn categories_decoder() -> decode.Decoder(List(String)) {
+  decode.one_of(decode.list(text_decoder()), [
+    text_decoder() |> decode.map(fn(cat) { [cat] }),
+  ])
+}
+
+fn cloud_decoder() -> decode.Decoder(Cloud) {
+  // Cloud element uses attributes: domain, port, path, registerProcedure, protocol
+  use domain <- decode.field("$attrs", decode.at(["domain"], decode.string))
+  use port <- decode.field("$attrs", decode.at(["port"], string_int_decoder()))
+  use path <- decode.field("$attrs", decode.at(["path"], decode.string))
+  use register_procedure <- decode.field(
+    "$attrs",
+    decode.at(["registerProcedure"], decode.string),
+  )
+  use protocol <- decode.field("$attrs", decode.at(["protocol"], decode.string))
+  decode.success(Cloud(domain:, port:, path:, register_procedure:, protocol:))
+}
+
+fn string_int_decoder() -> decode.Decoder(Int) {
+  use str <- decode.then(decode.string)
+  case int.parse(str) {
+    Ok(i) -> decode.success(i)
+    Error(_) -> decode.failure(0, "Int")
+  }
+}
+
+fn image_decoder() -> decode.Decoder(Image) {
+  use url <- decode.field("url", text_decoder())
+  use title <- decode.field("title", text_decoder())
+  use link <- decode.field("link", text_decoder())
+  use description <- decode.optional_field(
+    "description",
+    None,
+    decode.optional(text_decoder()),
+  )
+  use width <- decode.optional_field(
+    "width",
+    None,
+    decode.optional(int_text_decoder()),
+  )
+  use height <- decode.optional_field(
+    "height",
+    None,
+    decode.optional(int_text_decoder()),
+  )
+  decode.success(Image(url:, title:, link:, description:, width:, height:))
+}
+
+fn text_input_decoder() -> decode.Decoder(TextInput) {
+  use title <- decode.field("title", text_decoder())
+  use description <- decode.field("description", text_decoder())
+  use name <- decode.field("name", text_decoder())
+  use link <- decode.field("link", text_decoder())
+  decode.success(TextInput(title:, description:, name:, link:))
+}
+
+fn enclosure_decoder() -> decode.Decoder(Enclosure) {
+  // Enclosure element uses attributes: url, length, type
+  use url <- decode.field("$attrs", decode.at(["url"], decode.string))
+  use length <- decode.field(
+    "$attrs",
+    decode.at(["length"], string_int_decoder()),
+  )
+  use enclosure_type <- decode.field(
+    "$attrs",
+    decode.at(["type"], decode.string),
+  )
+  decode.success(Enclosure(url:, length:, enclosure_type:))
+}
+
+fn guid_decoder() -> decode.Decoder(#(String, Option(Bool))) {
+  use guid_text <- decode.field("$text", decode.string)
+  use is_permalink <- decode.optional_field(
+    "$attrs",
+    None,
+    decode.optional(decode.at(["isPermaLink"], bool_string_decoder())),
+  )
+  decode.success(#(guid_text, is_permalink))
+}
+
+fn bool_string_decoder() -> decode.Decoder(Bool) {
+  use s <- decode.then(decode.string)
+  case string.lowercase(s) {
+    "true" -> decode.success(True)
+    "false" -> decode.success(False)
+    _ -> decode.failure(False, "Bool")
+  }
+}
+
+fn skip_hours_decoder() -> decode.Decoder(List(Int)) {
+  use hours <- decode.optional_field(
+    "hour",
+    [],
+    decode.one_of(decode.list(int_text_decoder()), [
+      int_text_decoder() |> decode.map(fn(h) { [h] }),
+    ]),
+  )
+  decode.success(hours)
+}
+
+fn skip_days_decoder() -> decode.Decoder(List(Weekday)) {
+  use days <- decode.optional_field(
+    "day",
+    [],
+    decode.one_of(decode.list(weekday_decoder()), [
+      weekday_decoder() |> decode.map(fn(d) { [d] }),
+    ]),
+  )
+  decode.success(days)
+}
+
+fn weekday_decoder() -> decode.Decoder(Weekday) {
+  use day_str <- decode.then(text_decoder())
+  case string.lowercase(day_str) {
+    "monday" -> decode.success(Monday)
+    "tuesday" -> decode.success(Tuesday)
+    "wednesday" -> decode.success(Wednesday)
+    "thursday" -> decode.success(Thursday)
+    "friday" -> decode.success(Friday)
+    "saturday" -> decode.success(Saturday)
+    "sunday" -> decode.success(Sunday)
+    _ -> decode.failure(Monday, "Weekday")
+  }
+}
```

## src/webls/sitemap.gleam

Status: modified; +209; −0

```diff
@@ -1,9 +1,12 @@
+import gleam/dynamic/decode
 import gleam/float
+import gleam/int
 import gleam/list
 import gleam/option.{type Option, None, Some}
 import gleam/result
 import gleam/time/calendar
 import gleam/time/timestamp.{type Timestamp}
+import parsed_it/xml
 
 // Stringify ------------------------------------------------------------------
 
@@ -20,6 +23,34 @@
   <> "\n</urlset>"
 }
 
+/// Generates a sitemap index XML string from a sitemap index
+pub fn index_to_string(index: SitemapIndex) -> String {
+  let sitemap_content =
+    index.sitemaps
+    |> list.map(fn(ref) { ref |> sitemap_reference_to_string })
+    |> list.reduce(fn(acc, ref_string) { acc <> "\n" <> ref_string })
+    |> result.unwrap("")
+
+  "<?xml version=\"1.0\" encoding=\"UTF-8\"?>\n<sitemapindex xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\">\n"
+  <> sitemap_content
+  <> "\n</sitemapindex>"
+}
+
+fn sitemap_reference_to_string(ref: SitemapReference) -> String {
+  "<sitemap>\n"
+  <> "<loc>"
+  <> ref.loc
+  <> "</loc>\n"
+  <> case ref.last_modified {
+    Some(date) ->
+      "<lastmod>"
+      <> date |> timestamp.to_rfc3339(calendar.utc_offset)
+      <> "</lastmod>\n"
+    _ -> ""
+  }
+  <> "</sitemap>"
+}
+
 fn sitemap_item_to_string(item: SitemapItem) -> String {
   "<url>\n"
   <> "<loc>"
@@ -113,6 +144,40 @@
   SitemapItem(..item, last_modified: Some(modified))
 }
 
+/// Create an empty sitemap index
+pub fn sitemap_index() -> SitemapIndex {
+  SitemapIndex(sitemaps: [])
+}
+
+/// Adds a sitemap reference to the sitemap index
+pub fn with_index_sitemap(
+  index: SitemapIndex,
+  ref: SitemapReference,
+) -> SitemapIndex {
+  SitemapIndex(sitemaps: [ref, ..index.sitemaps])
+}
+
+/// Adds a list of sitemap references to the sitemap index
+pub fn with_index_sitemaps(
+  index: SitemapIndex,
+  refs: List(SitemapReference),
+) -> SitemapIndex {
+  SitemapIndex(sitemaps: list.flatten([index.sitemaps, refs]))
+}
+
+/// Create a sitemap reference with a URL location
+pub fn reference(loc: String) -> SitemapReference {
+  SitemapReference(loc: loc, last_modified: None)
+}
+
+/// Add a last modified time to a sitemap reference
+pub fn with_reference_last_modified(
+  ref: SitemapReference,
+  last_modified: Timestamp,
+) -> SitemapReference {
+  SitemapReference(..ref, last_modified: Some(last_modified))
+}
+
 // Types ----------------------------------------------------------------------
 
 /// A complete sitemap
@@ -127,6 +192,24 @@
   )
 }
 
+/// A sitemap index that references multiple sitemaps
+pub type SitemapIndex {
+  SitemapIndex(
+    /// The list of sitemap references
+    sitemaps: List(SitemapReference),
+  )
+}
+
+/// A reference to a sitemap within a sitemap index
+pub type SitemapReference {
+  SitemapReference(
+    /// The location URL of the sitemap
+    loc: String,
+    /// The time of last modification of the referenced sitemap
+    last_modified: Option(Timestamp),
+  )
+}
+
 /// A item within a sitemap
 pub type SitemapItem {
   SitemapItem(
@@ -152,3 +235,129 @@
   Yearly
   Never
 }
+
+// Decoders -------------------------------------------------------------------
+
+/// Result of parsing a sitemap XML - either a regular sitemap or an index
+pub type SitemapParseResult {
+  ParsedSitemap(Sitemap)
+  ParsedSitemapIndex(SitemapIndex)
+}
+
+/// Parses a sitemap XML string into a Sitemap
+pub fn from_string(sitemap_xml: String) -> Result(Sitemap, xml.XmlDecodeError) {
+  xml.parse(from: sitemap_xml, using: sitemap_decoder())
+}
+
+/// Parses a sitemap index XML string into a SitemapIndex
+pub fn index_from_string(
+  sitemap_xml: String,
+) -> Result(SitemapIndex, xml.XmlDecodeError) {
+  xml.parse(from: sitemap_xml, using: sitemap_index_decoder())
+}
+
+/// Parses a sitemap XML string, detecting whether it's a regular sitemap or index
+/// Returns a SitemapParseResult indicating which type was parsed
+pub fn parse(
+  sitemap_xml: String,
+) -> Result(SitemapParseResult, xml.XmlDecodeError) {
+  // Try parsing as regular sitemap first
+  case from_string(sitemap_xml) {
+    Ok(sitemap) -> Ok(ParsedSitemap(sitemap))
+    Error(_) ->
+      // Try parsing as sitemap index
+      case index_from_string(sitemap_xml) {
+        Ok(index) -> Ok(ParsedSitemapIndex(index))
+        Error(e) -> Error(e)
+      }
+  }
+}
+
+fn sitemap_decoder() -> decode.Decoder(Sitemap) {
+  // urlset is the root element, url children contain the items
+  // When there are multiple <url> elements, they become a list
+  // When there's a single <url> element, it's a single object
+  use items <- decode.field(
+    "url",
+    decode.one_of(decode.list(sitemap_item_decoder()), [
+      sitemap_item_decoder() |> decode.map(fn(item) { [item] }),
+    ]),
+  )
+  decode.success(Sitemap(url: "", last_modified: None, items:))
+}
+
+fn change_frequency_decoder() -> decode.Decoder(ChangeFrequency) {
+  use variant <- decode.then(decode.at(["$text"], decode.string))
+  case variant {
+    "always" -> decode.success(Always)
+    "hourly" -> decode.success(Hourly)
+    "daily" -> decode.success(Daily)
+    "weekly" -> decode.success(Weekly)
+    "monthly" -> decode.success(Monthly)
+    "yearly" -> decode.success(Yearly)
+    "never" -> decode.success(Never)
+    _ -> decode.failure(Never, "ChangeFrequency")
+  }
+}
+
+fn sitemap_item_decoder() -> decode.Decoder(SitemapItem) {
+  use loc <- decode.field("loc", decode.at(["$text"], decode.string))
+  use last_modified <- decode.optional_field(
+    "lastmod",
+    None,
+    decode.optional(timestamp_decoder()),
+  )
+  use change_frequency <- decode.optional_field(
+    "changefreq",
+    None,
+    decode.optional(change_frequency_decoder()),
+  )
+  use priority <- decode.optional_field(
+    "priority",
+    None,
+    decode.optional(decode.at(["$text"], string_float_decoder())),
+  )
+  decode.success(SitemapItem(loc:, last_modified:, change_frequency:, priority:))
+}
+
+fn string_float_decoder() -> decode.Decoder(Float) {
+  use str <- decode.then(decode.string)
+  case float.parse(str) {
+    Ok(f) -> decode.success(f)
+    Error(_) ->
+      // Try parsing as int and convert to float
+      case int.parse(str) {
+        Ok(i) -> decode.success(int.to_float(i))
+        Error(_) -> decode.failure(0.0, "Float")
+      }
+  }
+}
+
+fn timestamp_decoder() -> decode.Decoder(Timestamp) {
+  use date_str <- decode.then(decode.at(["$text"], decode.string))
+  case timestamp.parse_rfc3339(date_str) {
+    Ok(ts) -> decode.success(ts)
+    Error(_) -> decode.failure(timestamp.from_unix_seconds(0), "Timestamp")
+  }
+}
+
+fn sitemap_index_decoder() -> decode.Decoder(SitemapIndex) {
+  // sitemapindex is the root element, sitemap children contain the references
+  use sitemaps <- decode.field(
+    "sitemap",
+    decode.one_of(decode.list(sitemap_reference_decoder()), [
+      sitemap_reference_decoder() |> decode.map(fn(ref) { [ref] }),
+    ]),
+  )
+  decode.success(SitemapIndex(sitemaps:))
+}
+
+fn sitemap_reference_decoder() -> decode.Decoder(SitemapReference) {
+  use loc <- decode.field("loc", decode.at(["$text"], decode.string))
+  use last_modified <- decode.optional_field(
+    "lastmod",
+    None,
+    decode.optional(timestamp_decoder()),
+  )
+  decode.success(SitemapReference(loc:, last_modified:))
+}
```

## test/fixtures/robots.txt

Status: deleted; +0; −13

```diff
@@ -1,13 +1,0 @@
-Sitemap: https://example.com/sitemap.xml
-
-User-agent: googlebot
-Allow: /posts/
-Allow: /contact/
-Disallow: /admin/
-Disallow: /private/
-
-User-agent: bingbot
-Allow: /posts/
-Allow: /contact/
-Allow: /private/
-Disallow: /
```

## test/fixtures/robots/case_insensitive.txt

Status: added; +4; −0

```diff
@@ -1,0 +1,4 @@
+SITEMAP: https://example.com/sitemap.xml
+USER-AGENT: googlebot
+ALLOW: /posts/
+DISALLOW: /admin/
```

## test/fixtures/robots/flipped_order.txt

Status: added; +3; −0

```diff
@@ -1,0 +1,3 @@
+User-agent: googlebot
+Disallow: /admin/
+Allow: /posts/
```

## test/fixtures/robots/no_sitemap.txt

Status: added; +2; −0

```diff
@@ -1,0 +1,2 @@
+User-agent: googlebot
+Allow: /posts/
```

## test/fixtures/robots/robots.txt

Status: added; +13; −0

```diff
@@ -1,0 +1,13 @@
+Sitemap: https://example.com/sitemap.xml
+
+User-agent: googlebot
+Allow: /posts/
+Allow: /contact/
+Disallow: /admin/
+Disallow: /private/
+
+User-agent: bingbot
+Allow: /posts/
+Allow: /contact/
+Allow: /private/
+Disallow: /
```

## test/fixtures/robots/whitespace.txt

Status: added; +6; −0

```diff
@@ -1,0 +1,6 @@
+
+  Sitemap:   https://example.com/sitemap.xml  
+
+  User-agent:   *  
+  Allow:   /  
+
```

## test/fixtures/robots/with_comments.txt

Status: added; +6; −0

```diff
@@ -1,0 +1,6 @@
+# This is a robots.txt file with comments
+Sitemap: https://example.com/sitemap.xml
+
+User-agent: googlebot # Google's crawler
+Allow: /posts/
+Disallow: /admin/ # Keep admin private
```

## test/fixtures/rss.xml

Status: deleted; +0; −23

```diff
@@ -1,23 +1,0 @@
-<?xml version="1.0" encoding="UTF-8"?>
-<rss version="2.0.1">
-<channel>
-<title>Gleam RSS</title>
-<link>https://gleam.run</link>
-<description>A test RSS feed</description>
-<language>en</language>
-<category>Releases</category>
-<item>
-<title>Gleam 1.0</title>
-<description>Gleam 1.0 is here!</description>
-<link>https://gleam.run/blog/gleam-1.0</link>
-<pubDate>2024-08-11T20:22:50.481Z</pubDate>
-<guid isPermaLink="false">gleam 1.0</guid>
-</item>
-<item>
-<title>Gleam 0.10</title>
-<description>Gleam 0.10 is here!</description>
-<link>https://gleam.run/blog/gleam-0.10</link>
-<author>user@example.com</author>
-<guid isPermaLink="true">gleam 0.10</guid>
-</item></channel>
-</rss>
```

## test/fixtures/rss/full_channel.xml

Status: added; +27; −0

```diff
@@ -1,0 +1,27 @@
+<?xml version="1.0" encoding="UTF-8"?>
+<rss version="2.0.1">
+<channel>
+<title>Full Featured Feed</title>
+<link>https://example.com</link>
+<description>An RSS feed with all channel fields</description>
+<language>en-us</language>
+<copyright>Copyright 2024 Example Inc.</copyright>
+<managingEditor>editor@example.com</managingEditor>
+<webMaster>webmaster@example.com</webMaster>
+<pubDate>2024-01-15T12:00:00.000Z</pubDate>
+<lastBuildDate>2024-06-20T08:30:00.000Z</lastBuildDate>
+<category>Technology</category>
+<generator>webls</generator>
+<docs>https://www.rssboard.org/rss-2-0-1</docs>
+<ttl>30</ttl>
+<item>
+<title>Full Item</title>
+<description>An item with all fields</description>
+<link>https://example.com/full-item</link>
+<author>author@example.com</author>
+<comments>https://example.com/full-item/comments</comments>
+<pubDate>2024-06-20T08:00:00.000Z</pubDate>
+<source>Original Source</source>
+<guid isPermaLink="true">https://example.com/full-item</guid>
+</item></channel>
+</rss>
```

## test/fixtures/rss/minimal.xml

Status: added; +8; −0

```diff
@@ -1,0 +1,8 @@
+<?xml version="1.0" encoding="UTF-8"?>
+<rss version="2.0.1">
+<channel>
+<title>Minimal Feed</title>
+<link>https://example.com</link>
+<description>A minimal RSS feed</description>
+</channel>
+</rss>
```

## test/fixtures/rss/rss.xml

Status: added; +23; −0

```diff
@@ -1,0 +1,23 @@
+<?xml version="1.0" encoding="UTF-8"?>
+<rss version="2.0.1">
+<channel>
+<title>Gleam RSS</title>
+<link>https://gleam.run</link>
+<description>A test RSS feed</description>
+<language>en</language>
+<category>Releases</category>
+<item>
+<title>Gleam 1.0</title>
+<description>Gleam 1.0 is here!</description>
+<link>https://gleam.run/blog/gleam-1.0</link>
+<pubDate>2024-08-11T20:22:50.481Z</pubDate>
+<guid isPermaLink="false">gleam 1.0</guid>
+</item>
+<item>
+<title>Gleam 0.10</title>
+<description>Gleam 0.10 is here!</description>
+<link>https://gleam.run/blog/gleam-0.10</link>
+<author>user@example.com</author>
+<guid isPermaLink="true">gleam 0.10</guid>
+</item></channel>
+</rss>
```

## test/fixtures/rss/single_item.xml

Status: added; +11; −0

```diff
@@ -1,0 +1,11 @@
+<?xml version="1.0" encoding="UTF-8"?>
+<rss version="2.0.1">
+<channel>
+<title>Single Item Feed</title>
+<link>https://example.com</link>
+<description>A feed with just one item</description>
+<item>
+<title>Only Item</title>
+<description>The only item in this feed</description>
+</item></channel>
+</rss>
```

[Next page](https://git.cafe/versecafe/webls/pull/1/diffs?after=test%2Ffixtures%2Frss%2Fsingle_item.xml&expectedVersion=1&format=markdown)

This response is bounded; additional entries or patch content may be omitted.
