Original observation, one date, every page saved
On 2 September 2026, for 8 of 12 products in a frozen sample of Capterra category listings that we had classified as pure-play, at least one vendor-controlled page checked stated a numerical claim about assessment validity or predictive performance.
A validity citation study built from saved pages: for 12 products on a frozen roster we read the homepage, the science and research links in the header and footer, the sitemap pages whose path or title names validity, and the documents one click away, then coded every numerical claim about validity or predictive performance and traced each one to the source the page names.
Vendor-controlled assessment pages discuss validity, and some include numbers: a coefficient for job performance, a multiplier against interviews, a percentage of hires that exceed a goal. We wanted to know how many products on a fixed roster had a page that stated such a number, how many had a page that named a source for that number, and how many of those sources traced back to one well-known 1998 meta-analysis or its 2022 re-analysis. So we checked on one day, with every page saved.
This page reports what the checked pages stated and which sources they named, and makes no finding about the truth of any claim, the quality of any source or the reasons a page states what it states. SharpAssessment publishes this page and appears in no row.
The numbers
On 2 September 2026, for 8 of 12 products in a frozen sample of Capterra category listings that we had classified as pure-play, at least one vendor-controlled page checked stated a numerical claim about assessment validity or predictive performance.
For 6 of 12 products, at least one checked page mapped at least one such included numerical claim to an identifiable source; for 2 of 12, at least one checked page mapped at least one such included numerical claim to the Schmidt and Hunter (1998) lineage.
The product is the reporting unit; a claim may concern the product itself, a test family or a general construct such as structured interviews. Only included numerical validity claims count; customer-outcome statements that do not meet the inclusion rule and commercial outcomes are recorded in the dataset outside the count.
The secondary numbers, as prespecified: a source in the Sackett and colleagues (2022) lineage for 2 of 12; a vendor technical report for 2; a customer or case study named as the source of a primary claim for 2; another academic source for 1; another third-party source for 2; at least one included claim for which we found no identifiable source on the checked pages for 4; no included claim found on the checked pages for 4 (HR Avatar, HackerRank, Vervoe, Evalart); partially unresolved 0. Source kinds are not exclusive, so their sum can exceed 6.
How to cite this page
On 2 September 2026, for 8 of 12 products in a frozen sample of Capterra category listings that we had classified as pure-play, at least one vendor-controlled page checked stated a numerical claim about assessment validity or predictive performance. (SharpAssessment, Validity citation study, https://sharpassessment.com/assessment-validity-citations/).
The sentence is recomputed from the rows by a script; if the rows change on the quarterly recheck, the change log notes the date.
Method
Sample. The 12 products are the ones classified as pure-play in Capterra’s pre-employment testing category on 29 July 2026, the same frozen roster as the assessment pricing transparency index and the review cross-listing study. Products with few pages or no claim stay in; the optional cohorts 2 and 3 were not run.
What we read. We ran three routes on 2 September 2026 and recorded them per product. Route 1: the homepage and every same-domain page linked from its header, navigation or footer through an anchor whose text contains one of eight research or validity words listed in the codebook. Route 2 covered the vendor’s recursively expanded XML sitemaps; page-title requests returned usable responses for 14,836 URLs, while 848 translations were skipped and five returned 404 or 500. A page was included when its path or fetched title contained a validity word. Route 3 checked every document linked one click from those seed pages, keeping those whose final URL, content type or anchor text identified a research or technical document. A page whose own HTML declares a language other than English is excluded after fetch and listed with its declaration (six translated TestGorilla pages and the Evalart homepage). The sitemaps listed 15,689 URLs; 208 matched a path term, and the completed path-or-title pass added 232 entries to the manifest, yielding 226 distinct route 2 pages before the final language rule. Route 3 requested 2,270 linked destinations and kept 31. After redirect deduplication and the language rule, the checked set contains 266 pages: 11 homepages, 5 route 1 pages, 220 route 2 pages and 30 route 3 documents, each saved with its SHA-256 hash and UTC time.
Scan and coding. Every checked page was converted to text with table rows kept as rows. A unit became a candidate when it carried a coefficient-style number near a validity word, or any percentage, multiplier or fraction at all. Every embedded PDF figure of at least 300 by 300 pixels was processed with OCR row by row; a numeric row qualified when any row in the same figure or table contained a validity word. Every content image whose alt text, file name or nearby text contained a validity word was read, by eye when OCR could not read it. The scan read all 266 pages, 235 qualifying PDF figures (536 rows) and 202 image occurrences (158 files, 21 read by eye), and produced 990 candidate units. Each candidate has a stable ID built from the product, page URL, page hash and quotation and is classified in a ledger keyed by that ID as an included primary claim, customer-outcome claim, commercial-outcome claim or exclusion with a reason code; two published rules exclude the remainder, and each row records whether the decision was explicit or rule based. A passage that states more than one claim is split into one row per claim; a table or figure with one value per method is grouped once per method. Groups record the sources the page names, whether each is identifiable, and its lineage, which is 1998 or 2022 only when the page, or one documented hop from it, maps the number to the primary paper. The 990 units produced 999 claim rows, 348 with an explicit ledger decision and 651 excluded by the two rules. The 65 included rows form 48 text and PDF groups; 42 more groups are value rows from three figures, bringing the total to 90 primary claim groups; 17 rows are customer outcomes in 13 groups and 7 are commercial outcomes, and each of the 75 images containing a number has an explicit decision. The strict validation checks that no candidate containing a validity word, a figure label or an assessment term with an outcome term is classified by rule, that included rows and group occurrences coincide one to one, and that the saved files match their hashes. The protocol was corrected during six recorded rounds of independent review, each dated in the codebook.
The table
Twelve rows, alphabetical. “Primary claim groups on checked pages” counts distinct included claims after grouping; the source columns say whether at least one included claim maps to a source of that kind. “No included claim found on checked pages” describes the pages we checked, never what the vendor states anywhere else.
| Product | Pages checked | Primary claim groups on checked pages | Identifiable source | 1998 lineage | 2022 lineage | Vendor report | Case study | Other academic source | Other third-party source | At least one included claim for which we found no identifiable source on checked pages | No included claim found on checked pages | Partially unresolved |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Criteria | 47 | 5 | yes | no | no | no | yes | no | no | yes | no | no |
| Discovered (The Hire Talent) | 10 | 2 | no | no | no | no | no | no | no | yes | no | no |
| EmployTest | 6 | 1 | no | no | no | no | no | no | no | yes | no | no |
| eSkill | 15 | 6 | yes | no | yes | no | no | no | no | yes | no | no |
| Evalart | 6 | 0 | no | no | no | no | no | no | no | no | yes | no |
| HackerRank | 19 | 0 | no | no | no | no | no | no | no | no | yes | no |
| HR Avatar | 39 | 0 | no | no | no | no | no | no | no | no | yes | no |
| Prevue Assessments | 25 | 3 | yes | no | no | yes | yes | no | no | no | no | no |
| TestDome | 8 | 20 | yes | yes | no | no | no | no | no | no | no | no |
| TestGorilla | 70 | 46 | yes | yes | yes | yes | no | yes | yes | no | no | no |
| Vervoe | 17 | 0 | no | no | no | no | no | no | no | no | yes | no |
| Wonderlic Select | 4 | 7 | yes | no | no | no | no | no | yes | no | no | no |
What the pages stated
One block per product with an included claim; every group, occurrence, page hash and source is in the dataset.
Criteria (47 pages checked; 5 primary claim groups; 7 customer-outcome groups outside the count). On 2 September 2026 the page titled “What Are Pre-Employment Tests? | Criteria Corp” stated: “aptitude tests are 1.6x as predictive as unstructured job interviews”. The page names no source for this figure, and we found no identifiable source mapped to it on the pages checked. On 2 September 2026 the page titled “Case Study: Employment Tests Used to Predict Productivity | Criteria Corp” stated: “Additionally, it was determined that the validity coefficient between test scores and job performance was .26.” The page is the vendor’s own case study and reports the figure from the study it describes; we record a case study as the source. 3 more groups for this product are in the dataset.
Discovered (The Hire Talent) (10 pages checked; 2 primary claim groups). On 2 September 2026 the page titled “Assessment Validation - The Hire Talent” stated: “Through 30+ years of validating our assessments we are able to predict which employees will succeed on the job with up to 90% success.” The page describes thirty years of validation and a double-blind method in prose and names no document, so we found no identifiable source mapped to this claim on the pages checked. One more group for this product is in the dataset.
EmployTest (6 pages checked; 1 primary claim group). On 2 September 2026 the page titled “Pre Employment Assessment Test | Try for Free” stated: “Skills testing predicts job performance 5x better than interviews alone.” The page attributes the sentence to ResearchGate without naming a document, so we found no identifiable source mapped to this claim on the pages checked.
eSkill (15 pages checked; 6 primary claim groups; 1 customer-outcome group outside the count). On 2 September 2026 the page titled “Technical Skills Tests vs. Behavioral Assessments in Hiring | eSkill” stated: “To put it plainly: job-relevant technical tests (like work samples or knowledge assessments) are much better predictors of future job performance than personality tests, all ranking above 0.40 for validity coefficients.” The page names Sackett’s meta-analysis; we map it to the Sackett and colleagues (2022) lineage and record an academic source. 5 more groups for this product are in the dataset.
Prevue Assessments (25 pages checked; 3 primary claim groups). On 2 September 2026 the page titled “How Do You Prove Validity” stated: “You can see in the above-noted Technical Bulletin that the validity coefficients for the Prevue Assessments generally exceed .35.” The page points to a Technical Bulletin and to the Prevue Technical Manual; we record a vendor technical document as the source. 2 more groups for this product are in the dataset.
TestDome (8 pages checked; 20 primary claim groups). On 2 September 2026 the PDF document under testdome.com/evidence-based-hiring/ (ebh-standalone.pdf) stated: “Remember, the validity of unstructured interviews is 0.38, while the validity of structured interviews is 0.51.” The document identifies its source in footnote 34, which reads “Schmid and Hunter (1998)” and links to the Schmidt and Hunter paper; we map it to the 1998 lineage and record an academic source. The same document displays 18 validity coefficients in an image (Figure 4), from age at -0.01 to work-sample tests at 0.54; each row is one group mapped through the same footnote. 2 more groups for this product are in the dataset.
TestGorilla (70 pages checked; 46 primary claim groups; 4 customer-outcome groups outside the count). On 2 September 2026 the page titled “Revisiting the validity of different hiring tools: New insights into what works best - Part 2” stated: “For instance, the mean validity of .42 found for structured interviews by Sackett et al. (2022) is the average of the validity coefficient values reported by different studies included in their meta-analysis.” The page names Sackett et al. (2022); we map it to the 2022 lineage and record an academic source. On 2 September 2026 the page titled “Revisiting the validity of different hiring tools” displayed a figure, “Original and revised validity estimates for different hiring tools”, whose legend names Schmidt and Hunter (1998) and Sackett et al. (2022) and whose rows list, for example, structured interview .51 and .42 and cognitive ability .51 and .31; the ten rows are ten groups mapped to both lineages. The “Part 2” page also displays a second image, a 25-row table captioned “Adapted from Sackett et al. (2022)”; each row is a group in the 2022 lineage. 10 more groups for this product are in the dataset.
Wonderlic Select (4 pages checked; 7 primary claim groups). On 2 September 2026 the page titled “Data-Driven Candidate Research for Better Hiring - Wonderlic” displayed a graphic headed “THE MOST EFFECTIVE HIRING SELECTION PRACTICES”, credited to HBR.org, whose source line reads “SOURCE BASED ON DATA SHARED BY FRANK L SCHMIDT IN A NOV 6, 2013 ADDRESS TO PTCMW AS AN UPDATE TO: SCHMIDT, F. L. & HUNTER, J. E. (1998).”; its seven rows (for example cognitive ability tests .65, integrity tests .46) are seven groups recorded as a third-party source, not as the 1998 lineage, because the page attributes the graphic’s values to the 2013 address.
Scientific context
Two meta-analyses anchor most of the lineage findings. The Schmidt and Hunter (1998) paper tabulated validity estimates for selection methods: .54 for work samples and .51 for both general mental ability tests and structured interviews. The Sackett, Zhang, Berry and Lievens (2022) paper re-estimated many of the same methods with a different treatment of range restriction and reported lower figures for several procedures; its table lists work samples at .33, with a lower relative standing, and structured interviews at .42. In 2023 Oh, Le and Roth published a comment in the same journal (read here as the SSRN accepted-manuscript preprint) arguing that the 2022 paper’s case against correcting for range restriction in concurrent validation studies is not sufficiently supported and should be reconsidered; Sackett and colleagues replied that they endorse correction where applicant-pool information exists, that settings producing substantial restriction are uncommon among the studies feeding meta-analyses, and that their central conclusions hold. We have read all four texts in full; this page takes no side and treats neither the 1998 nor the 2022 estimates as settled or superseded. For what these figures mean when you evaluate an instrument for your own roles, see personality tests for hiring.
Limitations
One observation date. The three routes have bounded coverage: search-engine operators were unavailable to us, so the sitemap route is auditable but not claimed to be site-wide, and a claim on a page outside all three routes is not counted. A claim without a number, or with a number the published patterns do not catch, is outside the study. The codebook names the borderline coding decisions: a survey share of employers who report a benefit is not a validity claim, and a case-study outcome that states no criterion is a customer outcome. Nothing here assesses any product’s validity or advises which product to buy. Corrections: write to [email protected] with the row and the page you read; the observation date stays separate from the page’s modified date.
Dataset
Every number on this page comes from the frozen dataset, published in one folder under a SHA-256 manifest (version 2026-09-02, manifest hash 740f7dbf96728616, 31 files) with a README. The folder holds the CODEBOOK.md that defines every decision, the checked-pages-2026-09-02.json with their hashes and routes, the claims-2026-09-02.json and claim-groups-2026-09-02.json behind the evidence blocks, and the products-2026-09-02.json behind the table, together with the route, scan and image manifests, the decision ledger, the results file and the exact scripts. The saved pages, figures and images are not published because of their size; every row carries the SHA-256 of its file, and the files are available on request from [email protected].
Change log
- 2 September 2026: first observation, 12 products; headline 8 of 12, 6 of 12 and 2 of 12 as stated above.
- 2 September 2026, later the same day: the scientific-context paragraph now describes the 2023 comment by Oh, Le and Roth and the reply, after the full text of the comment was obtained; no number or row changed.
- Corrections to this observation, if any, will be listed here with a timestamp and the dataset version; later changes on vendor pages get a new dated observation.
- Next recheck due by 2 December 2026.