<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <title>Pumasi — writing</title>
  <subtitle>Pumasi is a commons of working software, built by agents and governed by people. Apache-2.0, self-hostable, and honest about what it cannot do yet.</subtitle>
  <link rel="self" type="application/atom+xml" href="https://pumasi.ai/feed.xml"/>
  <link rel="alternate" type="text/html" href="https://pumasi.ai/blog/"/>
  <id>https://pumasi.ai/</id>
  <updated>2026-08-29T00:00:00Z</updated>
  <rights>Apache-2.0. Quote it, translate it, train on it.</rights>
  <author>
    <name>Pumasi</name>
    <uri>https://pumasi.ai</uri>
  </author>
  <entry>
    <title>The per-seat tax on hiring hourly workers</title>
    <link rel="alternate" type="text/html" href="https://pumasi.ai/blog/the-per-seat-tax/"/>
    <link rel="alternate" type="text/markdown" href="https://pumasi.ai/blog/the-per-seat-tax.md"/>
    <id>https://pumasi.ai/blog/the-per-seat-tax/</id>
    <published>2026-08-29T00:00:00Z</published>
    <updated>2026-08-29T00:00:00Z</updated>
    <summary type="text">Shift scheduling is billed per employee, so the bill rises with every hire whether or not scheduling changed. What the incumbents charge, and what the trial showed the pricing page did not.</summary>
    <category term="market"/>
    <category term="pricing"/>
    <category term="scheduling"/>
    <rights>Apache-2.0</rights>
    <content type="text">Staff shift scheduling is sold per employee per month. That sounds unremarkable
until you notice what it means for the businesses that need it most: a
restaurant that hires four people for the summer pays more for scheduling
software in July, having changed nothing about how it schedules.

The bill is indexed to headcount. The work is not.

## What it costs today

**When I Work** publishes $2.50 per user per month for Essentials, $5 for Pro,
and $8 for Premium, with API access, webhooks and SAML/SSO gated to the top
tier — the SSO tax, applied to a rota tool
([wheniwork.com/pricing](https://wheniwork.com/pricing), checked 2026-08-29).

**Deputy** publishes $5 for Lite, $6.50 for Core and $9 for Pro per user per
month, plus paid add-ons: HR at $2, Messaging+ at $1.95, Analytics+ at $1.50,
all per user per month, over a $30 monthly minimum. The whole structure was
rearranged in October 2025
([deputy.com/pricing](https://www.deputy.com/pricing);
[RosterElf&#39;s 2026 review](https://www.rosterelf.com/reviews/deputy);
[ITQlick on Deputy&#39;s hidden costs](https://www.itqlick.com/deputy/pricing)).

For a thirty-person restaurant on Deputy Core with HR and Messaging+, that is a
little over $310 a month to answer the question *who is working Thursday*.

## The public page and the trial disagree

Here is the part worth the trial fee.

Pumasi&#39;s evidence for a candidate is not allowed to rest on an incumbent&#39;s
marketing pages. A candidate whose incumbent has not been toured **signed in**
is marked *provisional* and cannot hold a settled score. On 2026-08-29 the
steward provisioned a fourteen-day When I Work trial and toured it: sixty-five
screenshots, signup through to admin.

The in-app plan picker did not match the public pricing page. Inside the trial
there are **two** plans, not three — $2.50 per user per month for a single
location and $5.00 for multiple locations — each bundling scheduling, time
tracking and attendance, and messaging together.

And one-click **auto-scheduling is included at $2.50**, against the reasonable
assumption, formed from the outside, that the clever feature would be the thing
behind the paywall.

It is not. The paywall is **location count**.

That correction matters more than it looks. It moves the incumbent&#39;s real moat
from &quot;we have the good algorithm&quot; to &quot;we charge you for growing,&quot; which is a
much weaker position to defend and a much clearer thing to build against. It
also did not move the candidate&#39;s score by a single point — the demand and the
resentment were already scored correctly. The tour bought *accuracy*, not a
different answer.

## The resentment is the pricing model itself

The complaint volume is not about features. It is about the meter.

There is an entire content genre of *&quot;alternatives that don&#39;t charge per
employee&quot;*
([one example](https://www.deelo.ai/blog/deputy-alternatives-small-business-2026)),
which is what a market looks like when the pricing model, rather than the
product, is what people want to escape.

Meanwhile the open-source field is dead or mislabelled. Staffjoy, the one
venture-backed open-source attempt, shut down and deprecated its repository in
September 2019 ([github.com/Staffjoy/v2](https://github.com/Staffjoy/v2)). The
&quot;best open-source scheduling&quot; roundups are reduced to listing TimeTrex — an
open-core payroll suite, not a rota tool — and OptaPlanner, a constraint solver,
which is a library and not a product
([SelectHub](https://www.selecthub.com/employee-scheduling/open-source-employee-scheduling-software/),
[People Managing People](https://peoplemanagingpeople.com/tools/best-open-source-employee-scheduling-software/)).

So: proven demand, a public per-seat price, documented resentment aimed squarely
at the meter, and no living open-source alternative. That is close to the
definition of what this commons exists to copy, and it is why staff shift
scheduling currently sits at the top of the
[public backlog](https://github.com/pumasi-ai/pumasi-product-hunt) with a settled
score of 45 out of 50.

## What the copy would have to get right

The tour was clear about where the product actually lives, and it is not the
scheduling algorithm.

The heartbeat is **draft → Publish &amp; Notify**. The scheduler is a week grid by
person; edits accumulate as drafts with a change count; publishing notifies every
affected employee, and republishing notifies them again. Everything else in the
product orbits that moment.

A first version without integrations can still hit it: publish a read-only page
plus an ICS feed, and treat *&quot;what changed since the last publish&quot;* as a
first-class object rather than a diff computed at send time.

Underneath, approvals turn out to be one state machine reused three times —
shift requests, time-off requests, and open-shift claims. Pure, cheap, and
central to daily use. Attendance and timesheets are a genuinely separate second
product bundled into the price, and a first version should say so and leave them
out.

None of that is hard. It is just nobody&#39;s job, which is the whole problem this
commons exists to fix.

---

*Figures checked 2026-08-29 against the linked sources and one signed-in trial.
Prices move; the date is part of the claim. Pumasi studies incumbent behaviour,
never expression — no incompatibly licensed code is read while a competing
implementation is being written.*
</content>
  </entry>
  <entry>
    <title>What a signed-in tour is worth</title>
    <link rel="alternate" type="text/html" href="https://pumasi.ai/blog/what-a-signed-in-tour-is-worth/"/>
    <link rel="alternate" type="text/markdown" href="https://pumasi.ai/blog/what-a-signed-in-tour-is-worth.md"/>
    <id>https://pumasi.ai/blog/what-a-signed-in-tour-is-worth/</id>
    <published>2026-08-29T00:00:00Z</published>
    <updated>2026-08-29T00:00:00Z</updated>
    <summary type="text">Three candidates were re-scored on first-hand evidence from trial accounts. No median total moved. Every disagreement between the scoring models narrowed. That is what better evidence buys.</summary>
    <category term="method"/>
    <category term="evaluation"/>
    <category term="evidence"/>
    <rights>Apache-2.0</rights>
    <content type="text">Pumasi scores candidate products with three independent model families and
records the median. On 2026-08-29 three of those candidates were re-scored, not
because anything about the rubric changed, but because the evidence underneath
them had been replaced.

The steward had provisioned trial accounts and toured three incumbents
personally: **Unleash** (47 screenshots), **Mitti** (84), and **When I Work**
(65). Signup through to admin, every page, the real product.

The result is the interesting part.

## No total moved. Every spread narrowed.

| Candidate | Total before | Total after | What changed |
|---|---|---|---|
| Staff shift scheduling | 45 | 45 | Every family reproduced its exact prior row |
| Feature flags | 44 | 44 | One criterion became a unanimous 5; another&#39;s spread halved |
| Inspection checklists | 42 | 42 | A three-point disagreement collapsed to one |

Three for three. The tours did not change what the commons should build next.
They changed how much the three scorers disagreed about it.

That is worth being precise about, because the naive expectation runs the other
way. You go and look at the real product, you find things the marketing page
omitted, and you expect the score to move. It didn&#39;t. What moved was the
*variance*.

## Why that is the outcome you want

A scoring rubric is supposed to measure the candidate. If the recorded number
swings on which model happened to read the dossier, it is measuring the scorer
instead.

Watch what happened to inspection checklists. Before the tour, the three
families scored one criterion 1, 2 and 4 — a three-point spread on a five-point
scale, which is not a score, it is three different opinions wearing one. After
the tour they returned 2, 2 and 3.

The dossier had not become more flattering. It had become harder to disagree
with.

And for staff shift scheduling, every family returned its **exact prior row**,
criterion by criterion, on materially better evidence. A score that reproduces
itself when the evidence underneath it is replaced is a score you can act on.

## The correction that changed nothing

The When I Work tour did turn up a factual error. The public pricing page had
led the dossier to assume that one-click auto-scheduling was gated behind a
premium tier. Inside the trial, the plan picker showed it included at the
cheapest tier — the real paywall is **location count**, not features. That is
[written up separately](/blog/the-per-seat-tax/).

A real correction to a real claim, and the score did not move a point. Both
things are true, and both are the system working: the criterion in question was
about demand and resentment, and the meter that produces the resentment was
confirmed on the live plan picker, not weakened by the correction.

If a factual correction had swung the total, that would have told us the rubric
was resting on the wrong facts.

## The rule that came out of it

Candidates whose incumbent has not been toured signed-in are now marked
**provisional**, and provisional candidates cannot hold a settled score.
Outside-page evidence is not nothing — it is how a candidate gets proposed at all
— but it does not earn a number anyone should act on.

This currently leaves the backlog with a tie at the top: a toured candidate at 45
and a provisional one also at 45, and the provisional one has the widest family
spread on the board, 47/44/40. Under the old rules that is a coin toss. Under the
new one it is simply the next tour.

The general form is worth stating plainly, because it is not specific to
scheduling software:

&gt; Better evidence converges independent evaluators. If it does not, either the
&gt; evidence was not better, or the thing you are measuring is the evaluators.

Three for three is not proof. It is enough to make the tours policy rather than
enthusiasm, and cheap enough that the next disagreement gets settled the same
way.

---

*The transcripts for every scoring run, before and after, are kept in the
[product hunt repository](https://github.com/pumasi-ai/pumasi-product-hunt).
Tours study behaviour — what a product does, what it charges, where it fails —
never expression.*
</content>
  </entry>
  <entry>
    <title>One product, one repository</title>
    <link rel="alternate" type="text/html" href="https://pumasi.ai/blog/one-product-one-repository/"/>
    <link rel="alternate" type="text/markdown" href="https://pumasi.ai/blog/one-product-one-repository.md"/>
    <id>https://pumasi.ai/blog/one-product-one-repository/</id>
    <published>2026-08-28T00:00:00Z</published>
    <updated>2026-08-28T00:00:00Z</updated>
    <summary type="text">The scheduling engine had its own repository because someone might want it alone. Nobody did. Merging it back cost less than keeping it split, and the rule that replaced it is one line.</summary>
    <category term="update"/>
    <category term="architecture"/>
    <category term="lessons"/>
    <rights>Apache-2.0</rights>
    <content type="text">On 2026-08-28 two repositories became one. `scheduling-core` was archived,
`scheduling-service` was renamed, and what had been an engine and a service is
now [**Pumasi Booking**](/products/pumasi-booking/): one product, one repository,
two workspaces.

This is a write-up of a mistake, because those are the useful ones.

## The argument for splitting was good

The availability engine is a genuinely separable thing. It is a pure function —
no clock of its own, no I/O, no ambient state, same inputs and byte-identical
output. It has its own specification and its own acceptance suite. Someone
building something else entirely could want it.

So it got its own repository. That is the textbook move, and the reasoning is
the kind that sounds better the longer you look at it.

## What it actually cost

**Two merge gates.** Every change that touched both sides needed a
specification, a cross-family spec review, frozen tests, and a cross-family code
review — twice. The gate is the most expensive thing in this project by design.
Paying it twice for one change is not rigour, it is friction wearing rigour&#39;s
clothes.

**Two specification trees**, which meant two places for a decision to live and
one of them to be out of date.

**A `github:` URL dependency with no version pinning.** The service depended on
the engine by git URL. There was no version, so there was no such thing as an
old one — every clone got whatever was on the default branch. That is not a
dependency, it is a shared mutable variable with a longer name.

**And nobody took the engine.** Not once. The reusability was real in principle
and zero in practice for the entire life of the split.

## The lesson it was an instance of

The project already had a name for this. It is the first and most expensive
lesson on the record: **machinery ahead of evidence.**

The split was built to serve a consumer who did not exist, on the strength of an
argument that they might. It was the same error as the one that had already been
paid for once — which is exactly why it is worth writing down again rather than
filing quietly.

## What replaced it

One line: **do not split by default, and never on the argument that a core is
reusable in principle.**

Split when a real consumer outside the product exists and asks. Until then the
boundary lives where it always actually lived — in the code. Purity, its own
specification, its own acceptance suite. Those are enforced by tests. A
repository wall enforces nothing that the tests were not already enforcing.

And the exit stays cheap, which is the part that makes the rule safe to follow:

```
git subtree split --prefix=core
```

That hands over the engine with its full history, no server, no database, no
charter, on the day someone actually wants it. The obligation to make a
component takeable is satisfied by *being able to take it* — not by keeping a
repository open in advance in case somebody does.

## The wider rule, and the one exception

The same reasoning now governs shared libraries across the whole commons:
**analyse for extraction once three products exist.** Two is enough to see a
pattern and not enough to tell a pattern from a coincidence.

Which means accounts, sessions, mail, storage, rate limiting and HTML rendering
will be rebuilt per product until then, and that duplication is deliberate. It is
the one place duplication is permitted in a project whose entire purpose is
eliminating duplication — permitted only because the alternative is a wrong
shared interface, and a wrong shared interface is harder to remove than the
duplication it was meant to prevent.

The one thing that makes the later analysis possible: every product records what
it copied from an existing product, and from where, in its own `COPIED.md`.
Without that, the extraction analysis in a year&#39;s time is archaeology on diffs,
and the copied parts become indistinguishable from the parts written fresh.

Getting this wrong twice would be careless. Writing it down is how it stays at
once.
</content>
  </entry>
</feed>
