# `LangExtract`
[🔗](https://github.com/mdepolli/lang_extract/blob/v0.10.0/lib/lang_extract.ex#L1)

Extracts structured data from text with source grounding.
Maps extraction strings back to exact byte positions in source text.

This module is the main entry point: `new/2` builds a client,
`template!/2` builds a validated task definition (`template/2` is its
non-raising twin for runtime task data), `run/4` executes the full
pipeline, and `align/3` / `extract/3` expose the lower-level steps.
Beyond the facade:

  * `LangExtract.Prompt.Validator` — pre-flight check that few-shot
    examples align against their own source text
  * `LangExtract.Serializer` — convert results to plain maps and JSONL
    for storage or interop
  * `LangExtract.Extraction` — the extraction struct used in template
    examples and parsed LLM output

# `provider`

```elixir
@type provider() :: :claude | :openai | :gemini
```

# `align`

```elixir
@spec align(String.t(), [String.t()], keyword()) :: [LangExtract.Span.t()]
```

Aligns extraction strings to byte spans in source text.

Returns a list of `%LangExtract.Span{}` structs, one per extraction.

Designed for chunk-scale sources (the pipeline aligns against ~200-token
chunks). The fallthrough fuzzy phases scale super-linearly in source
tokens, so calling this directly on a book-length source can cost
seconds per unmatched extraction — for whole-document grounding, use
`run/4` or `stream/4`, which chunk first.

## Options

  * `:fuzzy_threshold` - minimum overlap ratio for fuzzy match (default `0.75`)

## Examples

    iex> LangExtract.align("the quick brown fox", ["quick brown"])
    [%LangExtract.Span{text: "quick brown", byte_start: 4, byte_end: 15, status: :exact}]

# `extract`

```elixir
@spec extract(String.t(), String.t(), keyword()) ::
  {:ok, [LangExtract.Span.t()]}
  | {:error, {:invalid_format, String.t()} | :missing_extractions}
```

Parses LLM output, aligns extractions against source text, and returns
enriched spans with class and attributes.

Accepts both canonical and dynamic-key format (where each entry uses
the class name as the key). JSON only (since 0.7.0). Strips markdown fences
and think tags before parsing.

## Options

  * `:fuzzy_threshold` - minimum overlap ratio for fuzzy match (default `0.75`)

## Examples

    iex> raw = ~s({"extractions": [{"class": "word", "text": "fox"}]})
    iex> {:ok, [span]} = LangExtract.extract("the quick brown fox", raw)
    iex> span.status
    :exact

# `new`

```elixir
@spec new(
  provider(),
  keyword()
) :: LangExtract.Client.t()
```

Creates a configured LLM client for extraction.

Raises `ArgumentError` on an unknown provider or unbuildable HTTP client
(e.g. missing API key). The raise is deliberate: misconfiguration here is
a programmer error caught at client construction, while runtime failures
during extraction stay data — per-chunk errors in `run/4`'s `Result`,
tagged tuples from `extract/3`.

## Examples

    client = LangExtract.new(:claude, api_key: "sk-...")
    client = LangExtract.new(:openai, api_key: "sk-...", model: "gpt-4o")
    client = LangExtract.new(:gemini, api_key: "gm-...")

# `run`

```elixir
@spec run(LangExtract.Client.t(), String.t(), LangExtract.Template.t(), keyword()) ::
  LangExtract.Result.t()
```

Runs the full extraction pipeline: prompt → LLM → parse → align.

Returns a `%Result{}` — document-ordered spans plus per-chunk errors.
This function cannot fail; `Result.errors` is the failure channel.
Every failure stays per-chunk: a chunk that fails to parse or times
out lands in `errors` as a `%ChunkError{}` with its byte range, and
the surviving chunks' spans are still returned. Chunk tasks run
linked, so a bug-level crash inside one propagates to the caller.

`LangExtract.Runner.run/4` shares this return contract, adding
retries, a shared request budget, and crash isolation (its supervised
tasks report crashes as `ChunkError`s too) on top. See the
failure-semantics table in the "Running in Production" guide.

## Options

  * `:max_chunk_chars` - chunk size in characters (default `1000`)
  * `:max_concurrency` - parallel chunk requests (default `10`)
  * `:task_timeout` - per-chunk task timeout (default `:infinity`)
  * `:fuzzy_threshold` - minimum LCS coverage for fuzzy match (default `0.75`)
  * `:min_density` - minimum matched-token density of a fuzzy span (default `1/3`)
  * `:accept_lesser` - allow prefix-fragment grounding as `:lesser` spans (default `true`)
  * `:exact_algorithm` - `:dp` (occurrence DP, default) or `:first_occurrence`

The chunk budget is measured in characters because it mirrors upstream
langextract's `max_char_buffer`: counting the same way keeps chunk
boundaries identical across the two libraries, which the cross-library
benchmarks depend on. Every output offset is bytes, and since chunking
is sentence-aware, boundaries can't be computed from the budget in
either unit — the byte ranges on results are the boundary source of
truth. See the "Alignment and Spans" guide for what the alignment
options tune.

## Examples

    client = LangExtract.new(:claude, api_key: "sk-...")
    template = LangExtract.template!("Extract entities.")

    %LangExtract.Result{spans: spans, errors: errors} =
      LangExtract.run(client, "the quick brown fox", template)

# `stream`

```elixir
@spec stream(LangExtract.Client.t(), String.t(), LangExtract.Template.t(), keyword()) ::
  Enumerable.t()
```

Streams per-chunk extraction results as each chunk completes.

Returns a lazy stream of `{:ok, %ChunkResult{}}` and
`{:error, %ChunkError{}}` events in **completion order**, not
document order — consumers who need latency don't wait for slow chunks;
consumers who need order sort by the byte ranges every event carries.
Nothing runs until the stream is consumed, and a slow consumer naturally
limits how many chunk requests are in flight.

Failure semantics match `run/4`: every failure stays per-chunk. A chunk
whose task times out is reported as
`{:error, %ChunkError{reason: {:task_exit, :timeout}}}` with its byte
range, and the remaining chunks keep flowing — `run/4` is exactly this
stream, collected and restored to document order.

Takes the same options as `run/4`.

## Examples

    client
    |> LangExtract.stream(document, template)
    |> Enum.each(fn
      {:ok, chunk_result} -> handle_spans(chunk_result.spans)
      {:error, chunk_error} -> log_failure(chunk_error)
    end)

# `template`

```elixir
@spec template(
  String.t(),
  keyword()
) :: {:ok, LangExtract.Template.t()} | {:error, Exception.t()}
```

Builds a validated extraction template, returning a tagged tuple.

The non-raising twin of `template!/2` for templates built from runtime
data (user-uploaded or JSON-loaded task definitions). Returns
`{:error, exception}` where `template!/2` would raise — an
`ArgumentError` for malformed example maps, or a
`LangExtract.Prompt.Validator.ValidationError` (carrying the per-example
issues) for examples whose extractions don't align.

## Examples

    iex> {:error, %ArgumentError{}} =
    ...>   LangExtract.template("Extract.", examples: [%{extractions: []}])

# `template!`

```elixir
@spec template!(
  String.t(),
  keyword()
) :: LangExtract.Template.t()
```

Builds a validated extraction template.

Examples are given as plain maps (string or atom keys, so JSON-loaded
task definitions work verbatim) or as ready-made structs. Map-authored
attributes are normalized to string keys, matching the wire format's
decoded shape; ready-made structs pass through unchanged. Each example's
extraction texts are validated against the example text using the
production aligner; misaligned examples raise
`LangExtract.Prompt.Validator.ValidationError` — a template that
constructs is a template whose examples align, unconditionally. For
runtime data where raising is inappropriate, `template/2` returns
tagged tuples instead.

## Examples

    iex> template =
    ...>   LangExtract.template!("Extract conditions.",
    ...>     examples: [
    ...>       %{text: "Patient has diabetes.",
    ...>         extractions: [%{class: "condition", text: "diabetes"}]}
    ...>     ]
    ...>   )
    iex> [example] = template.examples
    iex> example.extractions
    [%LangExtract.Extraction{class: "condition", text: "diabetes", attributes: %{}}]

---

*Consult [api-reference.md](api-reference.md) for complete listing*
