labmate logo: a chip with the letters l a b and gold pins

labmate

Local tools on Ollama: paper overviews and posts,
cited answers about a paper library, an arXiv scout and CV tailoring

GitHub

Why I built it

I wanted a hands-on way to learn the agentic patterns that matter in practice, on tasks I actually have: reading papers, writing about them, finding new ones, answering questions about my own research and tailoring a CV. labmate is a small collection of features that each show a different way of building with LLMs, all on one shared core.

Everything runs on local models through Ollama, with no paid APIs. The question each feature answers is who decides the next step: in a workflow the code fixes the order and the model fills in each step; in an agent the model picks its own tools. Both are built with LangChain chains, a LangGraph graph and create_agent. Model output is never trusted blindly: claims are checked against the source, and checks in code can send the model back or drop what it wrote.

Pipelines

Each feature is shown as the diagram it is built from. Hover a diagram to magnify it, click to open it full size.

plain code model call judge model agent check in code input / output

paper2flow chain

A research paper in, a fact-checked overview.pdf out: a cover, four cards (Task, Challenges, Method, Results) and data-flow diagrams.

  • Every claim carries a verbatim quote; a quote that is not in the paper, or has other numbers, is dropped in code.
  • A judge model checks each bullet; failed bullets are rewritten twice at most, then dropped.
  • The model only proposes typed boxes and arrows. Code validates them and writes the diagram.

paper2post chain

Reuses the paper2flow analysis to write a short LinkedIn post: a hook, three to five sentences that tell the story, and a question about something specific in the paper.

  • Every sentence cites known claims and passes the same fact-check as the overview.
  • Page 1 is the post text with icons and links, page 2 the pipeline diagram to attach as the image.

scout agent

Give it a topic and it searches arXiv, reads and writes notes.md. Nothing fixes the order of its steps: the model plans, picks a tool, reads what comes back and repeats.

  • Tools: search_arxiv, read_abstract, read_overview (runs paper2flow on one paper).
  • A citation guard refuses notes that cite an arXiv id the agent never read.

triage workflow + decision model

A topic and your interests in, a sorted reading list out: each arXiv hit becomes deep read, post or skip.

  • A small decision model scores each paper on five levels in a single pass, so one paper costs one short request.
  • Code turns the score into an action and keeps deep reads within a budget. With --run the chosen papers go through paper2flow or paper2post.
  • On 36 hand-labelled cases the default decision model picks the right action 30 times, a general 4B model 16.

ask: graph LangGraph graph

Questions about a library of PDFs (here my dissertation and papers), answered with citations such as Dissertation ยง4.2, p. 57. The code fixes the path.

  • Retrieval fuses BM25 and vectors; a judge grades the hits and the query is rewritten while evidence is missing.
  • An ambiguous question pauses the graph (interrupt) until you pick a reading.
  • A judge reads every answer sentence with its cited chunks; unsupported ones are dropped, and with nothing left the graph abstains instead of guessing.

ask: agent agent

The same task with the model in charge: it chooses what to search, whether to read more context and when to stop.

  • Tools: search_library, read_context, list_sources.
  • Its final text goes through the same citation parsing and sentence verification as the graph, so the two are comparable on a golden question set.

cv2job chain + one agent step

A CV (YAML) and a job posting in; a tailored CV, a cover letter and a separate gap report out.

  • Requirements are read from the posting and matched to CV bullets; every quote and id is checked.
  • The gaps step is an agent: it re-searches the CV with other words and may ask you a question.
  • Numbers and names must come from the source bullet, years of experience come from the CV, and the letter is a fixed template.

Shared core

Reply cache
  • A SQLite cache answers a repeated request (same messages, model, options and schema) without calling the model.
  • A rerun only pays for what changed.
Structured output
  • Replies are validated against a Pydantic schema, and a failed reply is retried with the error shown to the model.
  • The schema goes both into Ollama's format and into the prompt.
Grounding
  • Claims carry verbatim quotes that are checked against the source.
  • A judge-and-rewrite loop drops statements the evidence does not support.
Tracing & evals
  • LangSmith tracing is opt-in; nothing is sent otherwise.
  • Run metrics and a golden question set compare the ask graph with the ask agent, and a benchmark compares local models.

A personal, local-first project.


labmate is developed on a 32 GB Mac with Ollama and is not packaged on PyPI yet: you run it from a clone of the repository. The code and the diagrams above are the current state; expect fast-moving internals.


Stack: Python, LangChain, LangGraph, Ollama, SQLite, Typst, Mermaid. AGPL-3.0 licensed.

Star it on GitHub