Executive summary
Vivameda predicts which companies are structurally ready to enter hypergrowth. The top 1 percent of companies the model ranks experience hypergrowth at 27 times the base rate. The model was tested against 22 historical technology unicorns at their pre-scaling year, and 82 percent landed in the top 30 percent of scores. Of six unicorns held entirely out of training, 100 percent landed in the top 30 percent. The work is built on the Vivameda Longitudinal Workforce Panel: 4 million companies, 48 million observations, 1950 to 2020.
Research question
Most company scoring systems see growth only after it is already visible. The question is not whether growth is happening, but which companies are structurally ready to scale before the curve becomes obvious. The paper asks whether companies that later became unicorns looked different from their peers in the year before they scaled, using only data observable at that time.
The prediction target is hypergrowth, defined as year-over-year workforce growth above 100 percent within the next three years, rather than unicorn status. Valuation events are externally reported, delayed and unevenly captured. Workforce hypergrowth is observable directly in the panel and functions as a measurable scaling regime. The unicorn cohort is the validation showcase, not the training target.
Dataset and cohort
- 41,000 companies from the Vivameda panel in technology and technology-adjacent industries.
- 302,000 company-year observations across 2001 to 2016.
- Training set: 28,664 companies and 211,736 observations.
- Held-out test set: 12,285 companies and 90,233 observations, with 366 positive hypergrowth outcomes and a base rate of 0.41 percent.
- Validation cohort: 50 named technology unicorns, of which 22 had a visible pre-scaling year and 6 were held entirely out of training.
Methodology summary
The dataset is split by company rather than by year, so no company appears in both training and test sets, and company identifiers are not used as features. All features for a given observation are computed using only data available at or before year T, which prevents both time leakage and company-level memorisation.
Features include observed headcount, year-over-year growth, previous-year growth, growth acceleration, role diversity, role coverage, capability composition, capability coverage, primary role concentration, industry and company age. The model is a gradient boosted tree classifier with 200 trees, maximum depth 3 and learning rate 0.05, trained without class balancing. Four baselines are compared against the full model on the held-out test set.
Principal findings
1. The model predicts hypergrowth on held-out companies
The full model achieves AUC 0.889 on companies it never saw. At the top 1 percent of scores precision is 11.1 percent, a lift of 27 times the base rate. At top 5 percent precision is 4.5 percent, and at top 10 percent it is 2.8 percent. The top 1 percent of the test set contains 902 observations and 100 hypergrowth events, so reviewing 1 percent of the universe captures roughly 27 percent of all eventual scalers.
2. Workforce structure carries independent signal
A model using only capability composition, role diversity, role coverage, primary-role concentration and capability coverage, with no growth, size, industry or age data, achieves AUC 0.658 with a top 1 percent lift of 4.4 times and a top 10 percent lift of 2.6 times. Growth ranks companies; workforce structure explains whether an organisation is structurally ready for the move.
3. Historical unicorns concentrate in top scores
Of the 22 unicorns with a visible pre-scaling year, 36 percent landed in the top decile, 59 percent in the top 20 percent, 82 percent in the top 30 percent and 86 percent in the top 50 percent, against random expectations of 10, 20, 30 and 50 percent. For the six unicorns held out of training entirely, 50 percent landed in the top decile and 100 percent in the top 30 percent.
4. The signal is observable years before scaling
The prediction window is three years, and the unicorn validation scores companies at the year before their early-scaling pattern fires. The signature exists in workforce structure before headcount expansion becomes visible.
Commercial implications
For investors. Funnel compression and attention allocation. A reviewer screening only the top 30 percent of scores would have captured 82 percent of unicorns with visible pre-scaling years, and 100 percent of the held-out unicorns, against 30 percent for random review of the same volume. The model does not guarantee winners; it concentrates probability.
For accelerators and platforms. Application screening, founder benchmarking and programme selection. Companies in the top 1 percent enter hypergrowth at 27 times the base rate, while top 1 percent precision is 11 percent, so the correct operational mode is ranked review rather than threshold rejection.
For AI agents and go-to-market platforms. Account scoring, research automation and explainable company intelligence. The model logic and pre-computed signals are available as a knowledge file that an agent can ingest directly, so a buyer does not need to rebuild 70 years of workforce history.
Limitations
- Hypergrowth is the prediction target, not unicorn status. Hypergrowth precedes unicorn outcomes but is not identical to them, and the model is not a direct unicorn classifier.
- Workforce features overlap with growth in the full model. The strongest baseline of growth, industry and age slightly outperforms the full model on AUC. The commercial value of the workforce layer is enrichment and explanation, not replacement of growth-based scoring.
- Sample sizes for the unicorn validation are modest: 22 of 50 unicorns had a visible pre-scaling year and 6 were held out entirely, so estimates from the cleanest subset carry meaningful uncertainty.
- Selection effects and survivor bias. The cohort is selected for reaching unicorn status before 2024 with sufficient panel data, which biases it toward US technology-adjacent companies. Concentration in top deciles shows that unicorns shared the signature, not that all companies with the signature became unicorns.
- Workforce coverage is observed through public sources across 4 million companies and 70 years. Trajectories, growth rates and pattern firings are reliable as relative measures across companies and time periods.
Full paper and companion file
Published May 2026 by Vivameda Research. Free distribution permitted.
Discuss the research
Engagements start with a short data discussion call to identify the right starting point, whether that is the knowledge file, a custom scoring run on your own company list, a panel-scale signal licence, or full dataset access for model training.
- Contact the team
- Growth Intelligence — the layer behind the scaling signals
- Capability Intelligence — role mix and capability composition
- Compliance and methodology