# The method predictor

*How `methodpredictor/` turns a block count, a storey count, a floor plate and a
construction method into site mandays, a programme and a weekly headcount, before
there is a model, and which of its numbers are published, which are proposed and
which are this tool's.*

Every other manpower figure in this repository starts from an IFC take-off or
from the Annex C envelope over a known floor area. Early in a job neither exists:
there is a block count, a storey count, a floor plate and a choice of structural
and wall system. This page answers from those. It follows an input/output schema
proposed for "a predictor not based on the BIM model" (2026-09-25), with the
changes listed below.

---

## Inputs

| # | Input | Default | Notes |
|---|---|---|---|
| A1 | Project name | Block 123 Condo | label only |
| A2 | Blocks | 3 | |
| A3 | Typical levels per block | 30 | typical floors only |
| A4 | Floor plate, m² | 850 | per typical floor |
| A6 | Other floor area, m² | 0 | basement, podium, transfer, roof |
| A5 | Total construction area | *computed* | A2 × A3 × A4 + A6 |
| B1 | Structural method | Cast in-situ flat plate | Table 1 |
| B2 | Wall method | Dry partition | Table 2 |
| B3 | PPVC | No | Table 5 |
| C1 | Size the crew from | a target floor cycle | or a target total duration |
| C2a / C2b | Floor cycle, days / total duration, months | 6 / 30 | |
| C3 / C4 | Working days a week / hours a day | 6 / 8 | |
| C5 | Substructure, months | 6 | piling, pile caps, ground slab; more with basements |
| C6 | Fit-out after top-out, months | 12 | finishes, M&E testing, TOP |
| C7 | Stagger between blocks, weeks | 4 | 0 builds every block at once |
| C8 | Most workers the site can hold | none | stretches the programme to fit |
| C9 | Working days lost a year | 0 | public holidays, rain, MC — beyond the weekly rest day; the building types start from Singapore's 11 gazetted holidays |
| D1 / D2 | Terrain / site access | Flat / Good | Table 3 |
| D3 | Repetition bonus | from the level count | Table 4, or set by hand |
| — | Development type | Private residential | picks the Annex C baseline for the Code-anchored column |

## Starting from a site instead of a floor plate (A0)

Earlier than a floor plate there is a parcel, a Master Plan plot ratio and a
development type. Two multiplications cover the gap, and they are not the same
kind of number:

| Step | Formula | Basis |
|---|---|---|
| 1 | GFA = site area × gross plot ratio | **published** — URA Master Plan |
| 2 | construction floor area = GFA × asset-class multiplier | **assumed** — see below |
| 3 | typical = GFA · other = construction floor area − GFA | ePSS **definition** |
| 4 | floor plate = typical ÷ (blocks × levels) | arithmetic |

**The plot ratio is a ceiling, not a forecast.** It is the most the Master Plan
allows on the parcel. Site geometry, setbacks, height controls and the scheme
itself all take away from it, so step 1 gives the largest building that may be
built there and the page says so where the figure appears.

**The multiplier is this tool's assumption, and only the multiplier.** What
construction floor area *includes* is published: BCA's ePSS defines the Total
Floor Area a productivity figure is measured against as GFA plus environmental
deck, basement carpark, roof area including linkways, and other concrete area
including ancillary buildings. *How much* that adds for a given kind of
development is not published anywhere. `ASSET_CLASSES` in
`app/engine/landarea.js` carries a range per class with a reason for each; the
ordering — an industrial shed adds least, a hospital most — is the defensible
part. A reader who knows their own basement should type the multiplier, and it
is used as typed even past the class's range — a condominium with two basement
levels under most of its site is 1.4–1.7 — with a note that it is outside the
range. Only a figure below 1 or above `MAX_MULTIPLIER` (5) is clamped, and the
clamp is reported.

**The asset class is also the development type.** *Apply* sets the Annex C
category from the class (residential 2, commercial 3, industrial 4,
institutional 5), so the Code-anchored column reads the baseline for the
building actually chosen.

**Step 3 is why the split matters.** The predictor gives the typical floor area
the method index, the repetition bonus and the PPVC adjustment, and costs the
*other* floor area at the plain baseline rate with none of them. That is exactly
the line the ePSS definition draws: GFA is the floors that count towards plot
ratio — the repeating superstructure those adjustments describe — and everything
the multiplier adds is basement, carpark, roof and ancillary. Handing the whole
construction floor area through as typical area would apply a PPVC reduction to a
basement, which is the kind of error that makes a tool's answer better than the
truth.

## The calculation

The page runs it twice on the same inputs and prints both side by side:

* **As specified** — the schema as written: 0.6 m²/manday, and 70% / 30% of the
  typical floors' labour scaled by the structural / wall index.
* **Code-anchored** — the Code's own Annex C Table 4 baseline for the development
  type (0.319 private, 0.439 public, at the 2010 level), with the index applied
  only to the structural and wall trades: their share of a cast in-situ building
  in the trade mix, 33% and 15%.

The steps, each printed on the page with its formula:

1. **Baseline mandays** = typical floor area ÷ baseline rate.
2. **Method.** Structural part = baseline × structural share × 0.5 ÷ Ss. Wall part
   = baseline × wall share × 0.5 ÷ Sw. The rest is unchanged. 0.5 is the index of
   the conventional (two-way beam) building the rate describes.
3. **PPVC** cuts the site labour by 40% and puts 40% of the pre-cut labour in the
   factory. Without PPVC, the off-site labour is the structural part × Table 1's
   off-site share + the wall part × Table 2's.
4. **Repetition** takes Table 4's reduction off.
5. **Site conditions** multiply by terrain × access. The **other floor area** is
   added at the baseline rate with the same site factor, and no method, repetition
   or PPVC adjustment.
6. **Programme** = substructure + each block's floors at the cycle + (blocks − 1) ×
   stagger + fit-out. By total duration, the superstructure gets whatever the other
   phases leave; a target too short for them is lengthened, with a warning.
7. **Crew.** Person-days = site mandays × 8 ÷ hours a day. Each trade's labour is
   laid over its window of the programme as a trapezoid, and the weeks are summed:
   the peak is the busiest week. The structure gang per block is the structural
   trades' labour ÷ blocks ÷ each block's superstructure days.

### Who pays for which trade

The calculation decides how much labour each part of the building takes; the
trade mix only divides it. The structural part pays for formwork, rebar,
concrete, precast installation and the crane crew. The wall part pays for
masonry and plastering, which are named for the wall method (dry partition makes
them drywall partitions and skim coat). The rest pays for the finishing trades.
Piling and earthworks take 4% of the typical floors' labour, the whole-site
model's share. The other floor area is its own row. In the As-specified column
the schema's 70% is everything that is not walls, so it pays for the structural
and finishing trades together.

### The site limit

With C8 set, the page finds the smallest factor that stretches every duration
(cycle, both phases, the stagger) until the Code-anchored peak fits under the
limit. The labour does not change, only the weeks it is spread over. Both columns
then show that programme. The workbook carries the factor as an input. The
simulation fits the same factor once, on the central inputs, and every run
inherits it, so the panel and the result above answer the same constrained
question.

## Outputs

Peak, direct and subcontracted peak, average workforce, structure gang per block,
site and off-site mandays, floor cycle, total duration and site productivity for
both columns; the step-by-step trace; a weekly chart and table; the same weeks stacked by seven trade groups (`TRADE_GROUPS`, with a table view); a trade table
with each trade's weeks, mandays, direct and subcontracted split and peak; and
productivity against every Annex C category. Downloads:

* **Workbook (.xlsx)**: Summary, Inputs, Engine, Trades, Weekly, Shape, Tables.
  Every figure is a live formula over the Inputs and Tables sheets, with the
  value the page computed cached beside it. `npm run check-methodworkbook-excel`
  has Excel recalculate it under nine sets of inputs and compares every value
  with the engine.
* **Trace (CSV)**: every step, both columns.

The simulation adds three more, and they are **three files rather than one wider
one** because they are three different grains. A CSV is a table: a file whose
columns mean different things on different rows is readable by a person squinting
at it and loadable by nothing.

| File | Grain | Columns |
|---|---|---|
| `peak-workforce-samples.csv` | one row per iteration | `peak_workers` |
| `weekly-headcount-band.csv` | one row per calendar week | `week, p10_workers, p50_workers, p90_workers, coverage, sample_count` |
| `spread-attribution.csv` | one row per input per scope | `scope, week, input, input_label, rank_correlation` |

The peak-sample export is **unchanged** — widening it would have broken anything
already reading it, for no gain the other two files do not give.

Three things about the weekly export are deliberate:

- **The coverage cut-off is the chart's, exactly.** `profile.weeks` is already cut
  where coverage falls below the floor, and the file holds those weeks and no
  others, so the file and the chart describe the same programme. Exporting the
  weeks the chart declines to draw would hand a reader the least reliable figures
  the model produces, in a format that strips the caveat. A test asserts the file
  ends where the band ends and that no row survives below the floor.
- **`cutAt` and `longest` are not dropped**, they are stated in a comment line
  above the header — which every CSV reader skips and every person opening the
  file reads.
- **`sample_count` is there as well as `coverage`.** Coverage is a ratio, and a
  reader cannot recover the denominator from it: a week drawn from 501 runs and
  one drawn from 5,010 look identical at 50%. Anyone weighting these figures
  needs the count.

In the attribution file, `week` is empty for the `programme_peak` scope and
carries the week for the two week-scoped rows. Keeping it out of the weekly file
is what lets `week` mean one thing in every row of each.

Costs, at 10,000 runs: `weeklyProfileCsv` 1 ms for 4.1 KB, `attributionCsv` under
a millisecond. Both are byte-identical across two runs at the same seed.

## Checking against a finished job

Enter a finished job's inputs, then open its headcount: a CSV of `week` (or
`month`) and `workers`, optionally with a `project` column, or the estimator's
manpower-records template. It is drawn over the prediction, and the page compares
programme weeks, peak, average, site mandays and the productivity the record
implies. The file is read in the tab and sent nowhere.

## Quick estimate: the rule of thumb

The last section of the page is a calculator of its own, following a feasibility
proposal (2026-09-29, `app/engine/quickestimate.js`). It is kept because it is
simple, quick to use and easy to explain:

| Step | Formula |
|---|---|
| GFA | land area × gross plot ratio |
| CFA | GFA × asset multiplier (condominium 1.20, office 1.28, industrial 1.08, data centre 1.50), or the project's A5 total when no site is typed |
| Man-days | CFA ÷ productivity norm (0.40 cast in-situ, 0.55 precast, 0.85 DfMA/PPVC) |
| Average | man-days ÷ (months × 30.44 days × active work factor) |
| Peak | average × 1.5 to 2.0 |

Its fields start from the page:
- the site and plot ratio from A0;
- the multiplier and the active work factor from the building type (0.75,
  industrial 0.80, data centre 0.70);
- the norm from B1/B3;
- the duration from the detailed programme.

A figure typed over replaces its default, and clearing it restores the default,
so the section can be used alone. The proposal's worked example (12,000 m² × 2.8
× 1.20, 0.55, 24 months × 0.75, giving 134 average and 268 worst case) is
reproduced to the digit by the smoke test and by typing it into the page.

Below the steps, and closed by default, it is compared with the detailed answer.
The difference in peak splits into three ratios whose product is the whole of it:
the man-days, the working days they are spread over, and the peak's multiple of
the average. The largest of the three is named. None of its figures is published:
0.40 and 0.55 are the Code's private and public housing baselines raised by 25%, so
they describe development types, not technologies. The page says so.

## What the law requires on site

The trade mix sizes "crane and riggers" as a share of the labour, but the law
sizes a lifting team by cranes and work fronts. Below the headline, the page
counts what MOM requires, using the estimator's own rules from
`floorpredictor.js` so that the two tools cannot disagree:

* **While the frame goes up**, for each block lifting at once (frames overlap for
  ceil(superstructure days ÷ stagger) blocks):
  * one lifting supervisor per block (WSH (Operation of Cranes) Regulations 2011,
    reg 17);
  * for each crane (B4, default 1), a signalman (reg 19), riggers (reg 18: one
    compelled, two by practice) and a registered crane operator (reg 5).
* **On the whole worksite**, counted at the peak:
  * a WSH co-ordinator or registered WSH Officer (WSH (Construction)
    Regulations 2007 reg 6);
  * one first-aider per 100 people (WSH (First-Aid) Regulations 2006).

The lifting team is compared with the crane-and-riggers trade at its busiest
week. A shortfall is warned and shown with the peak it implies. It is not added
silently, because overriding one trade would stop the trades adding up to the
site. On the default three-block job, the law needs 15 and the mix gives 12. The
site appointments must be among the workforce (site management and trained
workers), so they are reported and never added.

## Building types

The tables describe one building, a repeating residential tower. A data centre,
an office and a warehouse use the same trades in very different proportions, so
the page offers six starting points (`app/engine/typologies.js`): public housing,
private condominium, commercial office, industrial or warehouse, data centre, and
school or hospital. Choosing one sets:

* **the Annex C category**, and so the published baseline, for the Code-anchored
  column. This is the only published part. Data centres sit on Business 2
  industrial land, so they take the industrial baseline (0.495).
* **the usual programme and method**: substructure and fit-out months, the usual
  structural and wall systems, and 11 days lost a year to public holidays.
* **a services intensity** (`servicesFactor`), which scales the labour of
  everything that is neither structure nor walls, and **trade scales**, which
  shift that labour between trades. For example, a data centre is ×2 with electrical
  and ACMV ×3, and a warehouse is ×0.8.
* **a list of what is different about building one**, in plain words.

Every value except the category is this tool's starting point for an expert to
replace, and the page says so. They are chosen to put the building types in a
defensible order, not to be right for any one job. The services intensity is
the figure most worth calibrating from a finished job's headcount.

## Changing the figures behind the answer

Every figure that is not a project input can be edited on the page: the method
indices, cycles and off-site shares, PPVC, repetition bands, site multipliers,
the loose constants and the whole trade mix. The registry is `ASSUMPTION_FIELDS`
in `methodpredictor.js`. Each field carries its range, its source and three
plain-words sentences: what it is, why it matters, and which way the answer
moves when it goes up.

Changes travel as a flat map on the inputs, `inputs.overrides`, for example
`{ 'structural.full_precast.lsi': 0.95 }`. A flat map can be counted, stored with
a scenario, compared between scenarios and written to the workbook, and an empty
map is exactly the published tables. The map has two layers: the building type's
and the expert's. The expert's layer wins and survives a change of building type.
Values out of range are clamped and unknown paths are ignored, and both are
reported rather than done silently. The workbook's Tables sheet holds the values
actually used, and its Summary lists every change. The Excel check replays a job
with expert tables and agrees with the engine.

The Code-anchored baseline can be replaced by hand (`codeBaselineRate`). The page
and the trace then say "set by hand, not Annex C", because it is the one figure
the page otherwise treats as published.

## Plain words, and what a change does to this project

**Explain every input** shows, under each field, what it is, why it matters and
which way the answer moves. It also gives the measured effect on the current
project: every numeric input raised 10% (a count by one), and every choice shown
with the peak it would give. Clicking a figure in the editor does the same for
that figure, or says why this project does not use it, for example a precast index
on a cast in-situ job. The sentence says which way the answer moves; the number
says how far for these inputs, and that differs between projects.

## Comparing scenarios

A scenario stores what a person chose: the form's inputs, the building type and
the expert's changes. It never stores the answers. Every comparison column is
recalculated from its scenario each time, so two scenarios differ only for reasons
they contain, never because the engine changed between the days they were saved.
Up to five are compared with what is on screen. For each one the page shows every
headline figure with its change from the current screen, the weekly headcount
lines on one axis, and a list of what differs. The figures a building type set
are counted in one line rather than listed trade by trade.

Scenarios stay in the browser and move between people as a JSON file
(`manpower-predictor-scenarios`, version 1). A file from a newer page still opens:
inputs and figures this page does not have are dropped and named.

The peak ÷ average ratio is shown beside a feasibility rule of thumb of
1.5–2.0× and flagged outside it. It is a second opinion only; the peak is still
the busiest week of the modelled curve.

## Where the numbers come from

| Numbers | Source |
|---|---|
| Annex C Table 4 baselines, the 25% open-solution target | **the Code** — read from `data/cop-buildability-2022.json` at load |
| Labour-saving indices, natural cycles, off-site shares (Tables 1–2), site multipliers (Table 3), repetition bands (Table 4), PPVC figures (Table 5), 0.6 m²/manday, 70 / 30 split | **the schema** — they follow the older Buildable Design Appraisal System; the 2022 Code publishes points, not indices |
| The 20-trade mix, direct / subcontracted shares, the 42-month phase template | **a second proposal** — its columns summed to 1.03, 0.96 and 0.92 and are only ever used as weights |
| Piling 4% | **the whole-site model** (`trademanpower.js`), itself assumed |
| Substructure 6 months, fit-out 12 months, stagger 4 weeks, trapezoid ramps of 20% | **this tool** — defaults for a Singapore high-rise job, every one an input |

None of the proposed or assumed numbers has been checked against measured site
returns. The headcount comparison above is how they get checked.

## What changed from the schema

1. PPVC's off-site labour is 40% of the labour **before** the site cut; the
   pseudocode took it after, which makes it 24%.
2. Tables 1–2's off-site shares are applied; the pseudocode never read them.
3. Steps 5–7 were cut off and are completed as above.
4. Table 4's last row, "40 18%", is read as more than 40 levels.
5. From the second proposal, three things were left out: a subcontractor
   multiplier of 1 ÷ (1 − 65%) on top of a crew that already counts everyone; a
   fixed peak of 1.4 × average; and a monthly profile with fixed crews that does
   not add up to the mandays.

## A check that it lands in the right place

A conventional building, two-way beam with brickwall, comes out in the
Code-anchored column within about 7% of the Code's own 0.319 m²/manday for private
housing. The smoke test holds it there.

## Open questions

1. **0.6 m²/manday.** Table 4 gives 0.439 public and 0.319 private at the 2010
   level. If 0.6 is a measured current figure, what is its source?
2. **What the structural index reaches.** Scaling 70% of the labour by the slab
   system also scales the M&E and the finishes, and gives 1.3 m²/manday on the
   schema's example, four times the Code's baseline.
3. **PPVC on top of B1/B2.** PPVC replaces the structural and wall systems, but its
   40% cut is applied on top of their indices. Should B1/B2 be ignored under PPVC?
4. **Other floor area.** It is built at the plain baseline rate. Should a podium
   take a method factor?
5. **The peak.** It runs at about 2½ times the average, because the piling months
   and the end of fit-out carry few people. Only measured headcount can say
   whether that is too peaky.

## Tests

* `node tools/smoke-methodpredictor.mjs` (in `npm test`): the schema's own example
  worked by hand, PPVC, both duration modes, the programme arithmetic, the site
  limit, the pools adding up for every method, the headcount reader, and the
  workbook read back.
* `npm run check-methodworkbook-excel`: the workbook in Excel against the engine.
* `npm run browser-check-predictor`: the page in a real browser.

---

## How sure is that? — the same model, run many times

Every input at concept stage is a guess inside a range. A single figure computed
from the middle of every range looks exactly like a measurement, and the question
a reader actually has — *how likely is it that we need more than 200 people* —
cannot be asked of it. `app/engine/montecarlo.js` draws each uncertain input from
a declared distribution and runs the chain again.

**It simulates the real model.** Each iteration calls `predict()` in full: the
method index, the repetition bonus, the PPVC split, the phase programme, the
trade trapezoids and the weekly summation. `predict()` costs about 0.12 ms, so
two thousand iterations is a quarter of a second and ten thousand is a little
over a second. A distribution of the real model beats a tight distribution of a
model nobody uses.

### There is no peak factor, and no active work factor

The usual early-stage shortcut divides mandays by duration for an average, then
multiplies by 1.3–1.6 for a peak. Both are already here, done from something:

- The **active work factor** is `daysPerWeek` and `hoursPerDay`, set against the
  BCA–MOM staggered rest day scheme, rather than one blanket 0.78 covering rest
  days, rain and holidays at once.
- The **peak factor** is not needed, because the shape is modelled. A static 1.45
  asserts every project's peak is the same multiple of its average. The weekly
  curve *derives* the ratio from where the trades sit, and it differs between a
  staggered three-block job and a single tower — `smoke-montecarlo.mjs` asserts
  exactly that, because if the ratio were constant this whole module would be an
  expensive way to reproduce a multiplication.

### Where A0's basis text comes from

`LAND_AREA_BASIS` in `app/engine/landarea.js` is the **single source of truth** for
what each A0 figure rests on. The page renders its basis lines, its explanatory
text and its two source links from that structure; `index.html` carries a
container and no prose.

It was not always so, and the reason for the change is the kind of drift this
repository exists to avoid. The markup used to restate the plot-ratio ceiling and
the ePSS definition in its own words, and carried **neither URL** — so a reader
could not reach the Master Plan or the ePSS FAQ from the tool that leans on them,
and the citation the tests assert was free to diverge from the claim on screen.

Each entry carries `kind`, `applies_to` (the A0 field), `agency`, `label`, `note`
and `url`. The link text is the **agency** — URA Master Plan, BCA ePSS — not the
address: a reader deciding whether to trust a figure wants to know who published
it, and a bare URL makes them hover to find out. Three checks hold this in place: a
unit test that every entry carries all six fields with a plausible agency and an
`https` link, a source-text test that the markup has not started restating it
again, and a browser assertion that every entry in the dataset reaches the page
with its own destination and `rel=noopener`.

### What is sampled, and what is not

**Never sampled: the Annex C Table 4 baseline.** It is published, exact for a
development category, and read from the Code dataset. Giving it error bars would
manufacture uncertainty its source does not have. Nor is the plot ratio: the
Master Plan states it. A test fails if any variation key names a rate.

Sampled: the declared assumptions, each with the distribution its shape deserves.
`DEFAULT_VARIATION` gives a reason per entry, and the ranges are *fractions* of
whatever the reader entered, so a 6-day cycle and a 12-day cycle vary alike.
Piling has the widest upward tail in the table, because the ground surprises a
programme in one direction.

### What is driving the spread

Rank correlation between each drawn input and the resulting peak — Spearman, not
Pearson, because a programme stretch that crosses a phase boundary moves the peak
in steps. This is what makes a distribution actionable: *"150 to 211"* tells a
reader to plan for a range; *"and most of that spread is the floor cycle"* tells
them what to go and pin down.

Two signs in that table are worth reading rather than skimming:

- **The floor cycle correlates negatively.** A slower cycle stretches the
  programme, and the same labour over more weeks is fewer people at once. A
  reader given only the magnitude would draw the opposite conclusion.
- **The substructure's sign flips with the sizing mode.** Under a floor cycle it
  is near zero; under a fixed total duration it is strongly positive, because
  piling overrun compresses everything after it into fewer weeks. One number
  could not have shown that.

### Week by week, not just at the peak

A peak is one number about one week; a site is staffed, housed and fed over the
whole programme. `predict()` already builds a weekly curve on every iteration, so
the band is free: per week, the P10, P50 and P90 headcount across the runs.

**Aligned by calendar week from site start, not by percentage of programme.** The
choice matters and is not the obvious one. Normalising every run to 0–100% would
compare *shapes*, which is a fair thing to do and not what a planner needs: they
hire, house and feed people in actual weeks. So week 90 means week 90.

The cost of that choice is that the later weeks are averaged over fewer runs,
because the runs have different lengths. A band drawn over three surviving
iterations looks exactly as solid as one drawn over two thousand, so:

- every week carries its **coverage** — the share of runs that got that far;
- the band is **cut** where coverage falls below `COVERAGE_FLOOR` (50%);
- the cut is **reported** — *"stops at week 123 of a longest run of 141"* —
  because a chart that simply ended early, with nothing saying why, would be its
  own kind of lie.

Coverage falling is also what proves the alignment: under a normalised axis every
run contributes to every bucket and coverage would be 1 throughout. A test
asserts it falls.

**The widest week is not the busiest week.** On the page's default inputs (2,000
runs, seed 42, Code-anchored) the band is widest in week 34 (50 to 279 workers)
and the median peak is week 56. The nominal substructure is six months, about
26 weeks, and week 34 is where the slower runs are still piling while the faster
ones have started the tower: it is where the runs *disagree* most, because a
piling overrun moves everything after it. The peak is where the site is
*busiest*.

They answer different questions, so the page names both and labels which question
each one answers:

| | Week 34 | Week 56 |
|---|---|---|
| What it is | the widest band | the highest median |
| Whose question | **programme and planning** | **resourcing and hiring** |
| What to do about it | settle the assumption driving it, and the range narrows | have the people hired, housed and fed *before* it |

`profile.widest` and `profile.busiest` carry those two weeks. `busiest` is the
week with the highest **median** headcount, which is deliberately not the same
statistic as the median of each run's own peak week: one asks *"in which calendar
week is the typical headcount highest"*, the other *"when did each run peak"*, and
averaging across weeks flattens a peak that moves. Both are reported rather than
one standing in for the other — the **week the peak falls in** is a distribution
too, because a peak in week 40 of a 116-week programme is a different hiring
problem from one in week 90.

### Which assumption is driving *that* week

The drivers table above attributes the **peak**, which is one number about the
whole programme. That is the right question for *"how many people at most"* and
the wrong one for *"why is week 34 so uncertain"*. So each named week carries its
own ranking: `sensitivityAt()` correlates the drawn multiples against that week's
headcount rather than against the peak.

It is worth doing because the answers differ. On the page's default inputs:

| Week | Top driver | Correlation |
|---|---|---|
| 34 — widest band | substructure and piling duration | −0.95 |
| 56 — highest median | floor cycle, days per storey | −0.64 |
| whole programme — peak | floor cycle, days per storey | −0.85 |

The widest week's ranking is led by piling; the busiest week's, and the
programme peak's, by the floor cycle. These are rank correlations, not shares of
the spread: −0.95 says the two move together almost monotonically, not that 95%
of the width is piling. A single table would have reported the average of
the two and pointed a reader at neither. A test asserts the two named weeks do not
blame the same input, because if they always did, two panels would be decoration.

**The alignment is the whole difficulty, and it is where this could have lied.**
Week 120's sample holds only the runs that lasted 120 weeks. Correlating it
against all 10,000 drawn multiples would pair one run's input with another run's
headcount and produce a number that looks like a finding. Each run's length is
recorded, and because the weekly samples are filled in iteration order, the runs
contributing to week *w* are exactly those that lasted longer than *w*, in that
same order. Where the reconstruction does not land on exactly that week's sample,
the function returns no drivers rather than a plausible-looking ranking built from
mismatched pairs. Two tests cover both halves of that.

A late week's sample is also *smaller*, so its correlations are noisier than an
early week's. That is the same coverage the band already reports, and it is said
on the page rather than left to be inferred.

An input a run does not read is not sampled and cannot appear as a driver:
the cycle or the total duration, whichever the sizing mode ignores, and the
block stagger on a single block, which the programme multiplies by
(blocks − 1).

### Two inputs that are deliberately absent from the driver list

A reader looking for *productivity* or *PPVC* in that table and not finding them
is entitled to know whether that means they do not matter. They matter a great
deal; they are absent for a reason, and the page says so rather than letting the
omission speak.

- **The baseline productivity rate is not sampled.** It is Annex C Table 4 of the
  Code — a published figure for the development category. Drawing it from a
  distribution would manufacture doubt about the one number on the page that is
  not in doubt. A test asserts it is never varied.
- **PPVC is not sampled.** It is a yes-or-no decision about how to build, not a
  quantity with a range, and it is already modelled *structurally*: it changes the
  floor cycle and cuts site labour by a fixed declared amount.

Both are answered by running the simulation twice and comparing, not by a
correlation. A correlation needs something continuous to correlate against, and
forcing either of these into one would have produced a number with no meaning.

### Reading the output

| For | Use |
|---|---|
| Dormitory and welfare sizing | P90 — a 10% chance of overflow |
| Canteen and amenity sizing | P50 — occasional overflow is tolerable |
| Risk conversation | the exceedance table: *"a 17% chance of needing more than 200"* |
| What to investigate next | the top row of the drivers table |
| Staffing and welfare over time | the weekly band, not the peak alone |
| Where a programme decision bites | the widest week, and its own driver list |
| How many to have hired, and by when | the busiest week — median for planning, P90 for headroom |
| Whether PPVC or a published rate matters | not the drivers table — run it twice and compare |

The deterministic single-value answer is printed beside the distribution. Where
it does not sit near the median, the ranges are lopsided — a floor cycle and a
piling programme both run late more readily than early — and seeing that is the
point of running both.

### What it is not

It is not a probability that the project will need *n* people. It is the spread
this model produces when its own declared assumptions are varied over their own
declared ranges. Every range is a judgement about a judgement, which the page
says. Narrow ranges would produce a confident distribution and no more truth.

## Calibrated against Singapore evidence (2026-09-29)

Research on 29 Sep 2026 checked the model against published Singapore practice.
Every source below was read on the cited page.

* **BCA's basis.** BCA measures site productivity (ePSS) as floor area completed per
  man-day **without site management**
  ([SCAL/SCCCI, Construction Productivity in Singapore, 2016](https://www.scal.com.sg/uploads/files/SCAL%20Guidebooks/Construction%20Productivity%20in%20Singapore%20-%20Copy%201.pdf)).
  The page reports `productivityBca` on that basis and compares it with Annex C. The
  earlier all-in figure is kept beside it.
* **Measured benchmarks.** Each is shown with its source:
  * HDB projects completed in 2021: 26.2% above the 2010 baseline of 0.439, which is
    0.554 ([HDB Annual Report 2021/22](https://assets.hdb.gov.sg/about-us/news-and-publications/annual-report/2022/developing-liveable.html));
  * the 2017 private non-landed industry average: 0.357;
  * PPVC at Clement Canopy: 0.613
    ([Wong et al. 2021](https://www.sciencepublishinggroup.com/article/10.11648/j.eas.20210606.12)).

  The feasibility proposal's "0.55 precast norm" is HDB's achieved figure.
* **Heat-stress rest (C10).** MOM requires 10–15 minutes of shaded rest an hour for
  heavy outdoor work at WBGT 32–33°C and above
  ([MOM, 2024](https://www.mom.gov.sg/newsroom/press-releases/2024/0906-revised-framework-to-guide-employers-and-protect-outdoor-workers-against-heat-stress)).
  The input takes the share of working time lost. It reduces the hours worked a day,
  so more people are needed for the same work; the programme is unchanged. It is also
  in the workbook, and the Excel check covers it.

## Public housing, fitted to recorded man-days (2026-09-30)

The predictor was run on the completed public-housing projects of a private record
set: each project's floor area, storeys, blocks, dates and PPVC, against the
man-days recorded for it. The records are not in this repository. Three figures
came out of it, and the public-housing building type now sets them
(`PUBLIC_HOUSING_FIT` in `app/engine/typologies.js`). They apply wherever that
building type is chosen: on this page and in the portfolio.

| Figure | Schema | Public housing | Why |
|---|---|---|---|
| Reference labour-saving index | 0.5 | **1.0** | The Code's public-housing baseline (0.439) was measured on HDB's own blocks, which were precast already. Against a reference of 0.5, precast halved the structural and wall labour a second time. |
| PPVC, site labour saved | 40% | **20%** | At 40% the PPVC projects came out about 15% below their records. |
| Other floor area, months on the template | 4–16 | **10–34** | A precinct's other floor area is a multi-storey car park, built beside the blocks and not ahead of them as a basement is. |

**What the first two move.** Predicted man-days against recorded:

| | Median project | All projects together |
|---|---|---|
| Schema's tables | 0.84× | 0.82× |
| With the fit | 1.02× | 0.98× |

Half the projects are within about 15% of their record and four in five within
about 35%. One project is not predicted to better than that: the fit removes the
bias, not the scatter.

On a precast project of four 16-storey blocks, site productivity on BCA's basis
goes from 0.573 to 0.480 m² a man-day, and with PPVC from 0.754 to 0.555.

**What the third moves.** The car park's man-days are unchanged. They now fall in
the same months as the frame and the early fit-out, so the peak is higher and
later: on the same project, 367 workers in week 101 instead of 276 in week 58. The
records hold no headcount by month, so this timing is an assumption and the peak is
not checked against them.

**Three things to know.**

* **It sits below HDB's published figure.** HDB reported 0.554 m² a man-day for
  projects completed in 2021. The fitted answer for a precast project is 0.480, and
  the benchmark table on the page shows the gap. The records run below HDB's
  figure in most years. The two may measure floor area or man-days differently;
  that is not settled here.
* **It reaches both columns.** The reference index is one table, so As specified
  is the schema's arithmetic on the fitted tables when public housing is chosen.
  With no building type, both columns are as before.
* **A less buildable system now costs more.** With precast as the reference, a
  two-way beam frame with brick walls comes out at 0.33 m² a man-day, near the
  Code's private baseline of 0.319.

**Not fitted.** Taller projects took a little more labour than predicted and the
lowest a little less, which is the opposite of the repetition bonus; the effect is
within 10% and the bands are left as the schema has them. Larger projects took
less labour a square metre than small ones. Neither is in the model.

### Reviewed on a larger set, and the split between disciplines (2026-10-06)

A second private record set covers 377 completed HDB projects, including all the
projects above. It records each project's on-site man-days split six ways:

* basement;
* structural;
* architectural;
* M&E;
* machinery and general workers;
* external works.

The six sum exactly to the total.

**The productivity fit holds.** The median project is 0.48 m² a man-day across all
377, the same as across the first set, so the three figures above are unchanged.

**The split between trades did not hold.** The shares are steady across sub-groups:
with and without a basement, and with and without PPVC.

| Discipline | Share of on-site man-days |
|---|---|
| Structural | about 35–39% |
| Architectural | about 27–30% |
| Machinery and general workers | about 15% |
| M&E | about 12–13% |
| Basement, where there is one | about 9% |
| External works | about 2% |

The proposal's trade mix gave a public-housing precinct M&E near 19% and general
labour near 10%. Public housing now scales the M&E trades by ⅔ and general labour
and site management by 2 (`PUBLIC_HOUSING_MIX`). The two factors move equal weight
in every column, so the total is unchanged and only who does the work moves. M&E
comes out near 12% and general workers near 15%.

Structural still reads below the record, about 29% against 35%. The calculation sizes the
structural part, not the mix, and that is open.

**The Buildable Design Score did not predict productivity.** Across the projects
that record both, the Buildable Design Score and the Constructability Score show no
relationship with on-site productivity: correlation about zero. One reason may be
that HDB projects' scores sit in a narrow band, about 80–88, so the records cannot show what a much lower or
higher score would do. But the scores here are compliance measures, and a higher one
should not be read as fewer man-days.

`node tools/compare-hdb-records.mjs` repeats both comparisons, the man-days and the
discipline split, for anyone who holds such records.
