The Big-Mad Behavioral Study · Draft of · not yet registered

The plan, before any data

Everything the study will test, how, and what would count as nothing, written down before anyone takes part. Once a registry timestamps it, it can be amended in the open but never quietly rewritten.

This is a draft, revised once after a methods review we asked an AI reviewer to do and say so. It goes to OSF Registries once the questions have been tried and the open decisions below are settled. Until then nothing here is a registration, and we will not call it one.

01

What it is

An observational study. Nobody is assigned to anything and nothing is changed in anyone’s work. For about a week, people report the moments that changed their mood or behaviour, as they happen, in their own words, by text or voice note.

The question is not how bad work feels, but where the frustration goes: onto the tool, onto other people, or back onto the self, and whether that is patterned by apps, dashboards, metrics or automation being part of the moment. Results are reported as patterns within people, never as how common something is in the population, and never as one moment causing another.

02

The hypotheses

H1, the main one

For the same person, a moment involving an automated or algorithmic system is more likely to land on people, someone else or themselves, than one of their moments that did not.

Against it: where the frustration goes has nothing to do with whether a system was involved. The difference has to be an odds ratio of at least 1.5. Smaller than that, an effect of this kind cannot be told apart from ordinary hour-to-hour confounds in a study this size, so we report it as no meaningful difference whatever its p-value.

H2, second

For the same person, those moments feel stronger: at least half a point higher on the 0 to 10 scale, about a quarter of a typical spread.

Exploratory

Whether H1 is stronger for people with more automation in their work. A test like this needs far more people than 60 to 90, so it is labelled exploratory everywhere it appears and cannot pass or fail the study. Comparisons between groups of people with more and less automation are described, never tested as the answer, because who signs up differs between those groups in ways we cannot control.

03

How each thing is measured

Whether a system was involved is read from what people wrote, never asked, because asking would put the idea in their heads. Two people code it independently from a written codebook, from transcripts, never audio, without knowing anything else about the person. They both code the same fifth of entries, and they must agree well (Cohen’s κ of at least 0.70) before the measure is used at all. If they do not, the main analysis is reported as failed rather than patched.

Where it went is the person’s own answer at each check-in: the tool, other people, myself, or nowhere, picking where most of it went. For the main analysis it is whether it landed on people (other people or myself) or not. How strong is one number, 0 to 10, with both ends described in every prompt. When is one tap: just now, within the hour, earlier today, or before today. And each person’s exposure to automation is a score from the screener, used as a number, not cut into groups.

04

Trying the questions first

Before the wording is fixed, 10 to 15 people from our own networks use the real text-message flow for two or three days, with a fifth choice, somewhere else, and space to say where. The four places are kept only if at least 90% of their answers fit them cleanly. Their answers are design data: never analysed, and none of them can join the study.

05

Who takes part, and when it stops

At least 20 people in each of three bands of automation exposure, 60 to 90 in all, in one wave. Unpaid, adults, in English for now, recruited in public, at first through LinkedIn and the people we know. Recruitment closes at 90 people or six weeks from launch, whichever comes first, and does not reopen to chase a result. The analysis begins only after the last person’s week ends, and runs once.

An unpaid public study draws people who already feel something about work and technology. That is the central property of this sample, not a footnote, and it is why the main question compares people with themselves. People will drop out, probably the most frustrated first; we report who finished against who started, and we do not fill in missing answers.

06

The analysis

H1 is a multilevel logistic regression, entries within people: person_directed ~ system_involved + (1 | participant). We try letting the effect vary by person, and if that model will not fit, as is likely with five to seven entries each, we fall back to the simpler one. That fallback is declared here so it cannot be chosen after seeing the data.

Only people with at least one moment of each kind can inform the comparison. If fewer than 30 do, the main analysis is reported as underpowered and descriptive only. Three checks are declared in advance: the same model on moments reported within the hour, on one entry per day, and with time of day added. H2 is the same comparison on intensity. Moments reported as before today are left out of every confirmatory test and counted separately. Every subgroup, by job, channel or anything else, is exploratory and labelled so.

07

Ethics

An independent review board is to review the study before it opens. Consent comes before any data, and says plainly that the study is unpaid. Voice recordings get their own written permission, are transcribed on our own computers and never go to an outside company. A plan for anyone in distress, with crisis lines, is in place before the first check-in.

08

Still to decide

Who runs the models (they will not code), which recruitment channels beyond the first, the review board and its reference, and what, if any, de-identified data is shared, which needs a check for re-identification first. Each blocks registration until it is settled, and each will be settled here, in public.

The limits are written out on the study’s page, and the consent people would be asked for is here.