HomeAboutServicesMethodsHow we workProgrammesIntegrityTeamInsightsContactSign in

Research data sourcing and analysis

Your data has to survive a reviewer who is looking for a reason to say no.

NextDataLab is the research data arm of NextPUBGlobal. We design the study, source the data, run the analysis and write the methods — so that when a referee questions your sampling, your assumptions or your model, there is an answer already documented.

A vertical of NextPUBGlobal · A Careerantra Solutions company

PROJECT ND-4417Cross-sectional · 6 sitesLIVE

Every project runs the same nine stages, and you can see which one you are in. Illustrative view of the client dashboard.

Reporting and integrity standards we work to

CONSORTPRISMA 2020STROBE COREQSRQRCHEERS ARRIVEFAIR data principles COPE guidanceICMJE authorship criteria
0Stages from research question to archived dataset
0Analytical methods in the open library
0Statistical and qualitative environments in use
0Independent checks before any result leaves us

Where papers actually fail

Rejection is usually a methods decision made months earlier.

A manuscript that reads badly can be rewritten in a fortnight. A study whose sample cannot carry its claim cannot be rewritten at all. By the time a referee says so, the funding is spent, the cohort has dispersed and the only honest fix is to collect again.

The three sentences that end a paper are almost always the same ones. The sample size is not justified. The analysis does not match the design. The results cannot be reproduced from what is described. None of them is about writing. All of them are decided before a single word of the manuscript exists.

NextDataLab exists to take those decisions early, on the record, with someone accountable for them. That is the whole proposition: not a faster analysis, but an analysis that was chosen for a reason you can state out loud.

Why we were set up

What we do

Eleven points of support, in the order a study actually moves through them.

Most researchers come to us in the middle — with data already collected, or an analysis a reviewer has questioned. We take work at any point. It costs less when it starts at the top.

Design and measurement

Protocols, variable maps, sampling frames, a priori power calculations, instrument development and validation, ethics and consent documentation.

  • Research design and protocol
  • Sampling and power analysis
  • Instrument development and validation
  • Ethics and consent documentation

Getting the data

Fieldwork we run or supervise, and existing data located, licensed and merged into something you can actually model.

  • Primary data collection
  • Secondary data sourcing
  • Data cleaning and preparation

Analysis and reporting

The analysis the design supports, run correctly, checked independently, and written up so a referee can follow every number back to a line in a script.

  • Quantitative analysis
  • Qualitative analysis
  • Evidence synthesis and bibliometrics
  • Reporting, reproducibility and reviewer response

Our promise

We change how the analysis is done. We never change what it found.

We will tell you which test your design supports, run it correctly, and write the methods so a referee can follow them. What we will not do is move a threshold, drop a case that spoils a distribution, or report the one specification of twelve that worked.

Read the full integrity position

Findings

Reported as computed, including the null ones.

Specifications

All disclosed, not only the best.

Your raw data

Never altered, always returned.

Authorship

Yours. We go in the acknowledgements.

Disciplines

Research we support.

Subject specialists across the fields our clients publish in, including interdisciplinary work that does not sit neatly inside one of them.

Medicine and health sciencesClinical trials, epidemiology, public health, nursing, dentistry, pharmacy
Management and commerceOrganisational behaviour, marketing, finance, operations, entrepreneurship
Social sciencesSociology, psychology, political science, anthropology, social work
Economics and policyDevelopment, labour, health economics, impact evaluation, econometrics
EducationPedagogy, assessment, higher education studies, learning analytics
Engineering and technologyIndustrial engineering, computing, energy systems, quality and reliability
Agriculture and environmentAgronomy, food science, climate, sustainability, natural resources
Life sciencesBiotechnology, microbiology, genetics, ecology, veterinary science
Law and humanitiesEmpirical legal studies, media, linguistics, mixed-methods humanities work

How we work

Two programmes. One method.

For the individual researcher with a study to finish, and for the institution that needs research capacity it does not have on staff. Both run the same nine stages, with every check placed before the point where fixing it becomes expensive.

For researchers and scholars

Decide the analysis before you collect.

Choosing a test after seeing the data is how p-hacking happens by accident. In our sequence the analysis plan is Stage 3 and collection is Stage 5. Fixing the plan first is what makes the eventual result defensible — and it is free to do.

  • Methodology consultation, from USD 120
  • Single analysis, fixed scope and fee, from USD 450
  • End-to-end study, design through deposit
  • Reviewer response and reanalysis, from USD 300

The nine stages

For institutions and departments

Shared analysis capacity, on tap.

A statistics and data unit your faculty can book, without the cost of hiring one — plus the training that reduces demand on it over time, because a department that can size its own samples stops producing unpublishable studies.

  • Named analysts allocated to your institution
  • Workshops on design, sampling, software and PRISMA
  • Research data management policy and FAIR setup
  • Quarterly reporting on usage and research output

Institutional programme

Insights

Notes from the analysis desk.

Practical, source-checked writing on research data and statistics — the things nobody explains until a reviewer has already asked.

Sample size

The power calculation your reviewer will ask for, and when it is too late to run it

Why a post-hoc power analysis convinces nobody, what an a priori calculation needs as input, and how to report it in three sentences.

Read the note
Assumptions

Your data violated normality. That is usually not the problem you think it is

When the assumption actually matters, when the central limit theorem covers you, and the three alternatives worth considering before you transform anything.

Read the note
Missing data

Listwise deletion is a decision, not a default

How to state a missingness mechanism, when multiple imputation is worth the complexity, and what to write in the methods either way.

Read the note

All notes

Start with a free methodology review.

Send the proposal, the dataset description or the reviewer's comment. We come back with what the design supports, which analysis fits, what it would cost — or an honest note that you do not need us.

Send your brief Meet the team

About NextDataLab

The publishing problem usually begins as a data problem.

NextPUBGlobal spends its days on whether a manuscript will survive an editor. Working on that, one pattern kept repeating: the papers that failed were rarely badly written. They were papers where the sample could not support the claim, the analysis did not match the design, or nobody could reconstruct how a number was produced. NextDataLab exists to fix that earlier, where it is still cheap.

2026Established as the research data vertical of NextPUBGlobal
9Disciplinary desks, from clinical trials to empirical legal studies
0Papers we have accepted authorship on, by policy

Why we exist

Editing cannot rescue a study that was designed wrong.

Language editing is a real service and it solves a real problem. But it operates at the end of a process, on a document, after every consequential decision has already been made. A copy editor can fix a sentence about a sample. Nobody can fix the sample.

Sitting inside a publication support business, that limit becomes obvious quickly. You watch a well-written manuscript come back from a journal with three statistical comments, and you watch the author discover that answering them properly means going back eighteen months. The paper was not rejected because of how it read. It was rejected because of a decision taken before the first participant was recruited.

NextDataLab was set up to move upstream of that moment — into design, sampling, measurement, collection and analysis — and to be accountable for the decisions taken there. The unit of work is not a document. It is a study.

A null finding you can defend is worth more than a significant one you cannot.

What changes when the data comes first

  • The sample size has a written justification before recruitment, not a post-hoc rationalisation after
  • The analysis is pre-specified, so choosing a test is a design decision rather than a reaction to the data
  • Assumptions are tested and reported instead of assumed and hoped for
  • Every number in the manuscript traces to a line in a script that a stranger can run
  • The methods section is written from the protocol, not reverse-engineered from the output
  • A statistical reviewer's comment is answered with a reanalysis, not an argument

What does not change

  • The research question is yours
  • The interpretation is yours
  • The authorship is yours
  • The result is whatever the data says it is

Who we are for

Three questions worth answering before you write to us.

Whether we are the right people, whether this is the right moment, and what physically arrives at the end.

We work with

  • PhD scholars and postdoctoral researchers
  • Faculty and principal investigators
  • Clinicians running audits and trials
  • University research cells and IQAC teams
  • Policy institutes and think tanks
  • Funded projects with a data workstream

We are brought in when

  • The proposal needs a defensible sample size
  • Data exists but nobody can analyse it
  • A reviewer has challenged the statistics
  • A thesis committee wants the methods tightened
  • A dataset must be prepared for deposit
  • A department needs analysis capacity it lacks

What you get back

  • A cleaned, documented, version-controlled dataset
  • Scripts that reproduce every number in the paper
  • Journal-ready tables and figures
  • Methods and results text you can defend
  • A codebook and data availability statement
  • A named analyst who answers the reviewer

How we operate

Six commitments that decide what we will take on.

These are not aspirations. Each one has cost us work, and each one is the reason a supervisor or an ethics committee can accept our involvement without a conversation.

The design comes before the data

Where you engage us early enough, the analysis plan and the sample size are fixed and written down before collection begins. Where you do not, we say plainly what the existing data can and cannot support.

Every number is traceable

Nothing is done by hand in a spreadsheet. Cleaning, recoding, exclusion and analysis all live in a script with a changelog, so any result can be traced back to the raw file it came from.

Two analysts, never one

The person who runs the model is not the person who reproduces it, and neither of them signs off the methods text. Three separate checks, deliberately not performed by the same pair of hands.

We do not write your paper

We write methods and results — the parts that describe what was done and what was found. Introduction, discussion and conclusion are yours, because they are the parts that constitute authorship.

We tell you when you do not need us

The first scoping call is free and ends with a recommendation, not a quote. Where a supervisor or a departmental statistician can answer the question, that is what you will hear.

Your data leaves as it arrived

Raw files are never altered. They are returned, and deleted from our systems at the end of the agreed retention period, with the deletion documented.

The order of work

What a study looks like from our side.

Compressed here to five moments; the full sequence runs to nine stages, each with its own deliverable.

STAGE 01Someone sends us a question

A proposal, a dataset description, or three sentences from a referee. We read it and reply with what the design can carry.

STAGE 02 – 04The study is fixed on paper

Variables, hypotheses, sampling frame, power calculation, instrument, ethics pack. Everything that is expensive to change later is decided here.

STAGE 05 – 06Data is collected or sourced, then cleaned

Fieldwork we run or supervise, or existing datasets located and merged. Cleaning happens in a script, with the missingness mechanism stated.

STAGE 07 – 08The pre-specified analysis is run, then re-run

By us, then independently from the raw file by an analyst who was not on the project. If a number does not reproduce, nothing ships.

STAGE 09It is written up and archived

Methods and results text, journal-ready tables, the reproducibility pack, the repository deposit and the data availability statement.

Open every stage in full

Part of NextPUBGlobal

NextDataLab is the research data vertical of NextPUBGlobal, a Careerantra Solutions company. The publishing side works on manuscript development, journal selection and editorial compliance; this side works on everything that has to be true before a manuscript is worth writing.

Clients use one, the other, or both. Where both are involved, the data team hands over a methods section the editorial team does not have to invent, and the editorial team flags reviewer comments the data team answers with a reanalysis.

The two are deliberately staffed and checked separately. An editor does not sign off statistics, and an analyst does not sign off prose.

Where we are

Unit 603–604, Floor 6, Tower B, Bhutani Alphathum, Sector 90, Noida 201305, India. We work with researchers and institutions internationally, in English, across time zones, and bill in USD, EUR, GBP or INR.

Contact the desk

If the study is still on paper, this is the cheapest hour you will spend on it.

A scoping call costs nothing and ends with a written note on what the design supports — including, sometimes, that you do not need us.

Book a scoping call See the services

What we do

Eleven points of support, in the order a study actually moves through them.

Most researchers come to us in the middle — with data already collected, or an analysis a reviewer has questioned. We take work at any point on this list. It costs less, and the result is stronger, when it starts at the top.

01 – 04Design, sampling, measurement and ethics, before anything is collected
05 – 07Primary collection, secondary sourcing, cleaning and codebook
08 – 11Analysis, synthesis, reporting and reviewer response

How to read this list.

The eleven entries below are sequential, not a menu of alternatives. Each one assumes the decisions taken in the one above it, which is why entering at Stage 8 with an undocumented dataset is more work than entering at Stage 2 with nothing but a question.

Each entry says what the work involves and what physically arrives at the end of it. Where you already have that deliverable — a validated instrument, an approved ethics pack, a clean dataset — we start after it and price accordingly.

Where people usually enter

  • Proposal or synopsis stage — Stage 1, the cheapest and most useful place to start
  • Data collected, analysis not begun — Stage 7, our most common entry point
  • Analysis done, committee unconvinced — Stage 8, usually a re-specification
  • Statistical comments from a journal — Stage 11, with a reanalysis and a response letter
  • Dataset to be deposited — Stage 11 alone, for a funder or repository deadline
01

Research design and protocol

Turning a question into a testable design: variables, hypotheses, unit of analysis, comparison groups, and the reporting guideline the study will be written against. The output is a protocol you can hand to a committee, not a conversation you have to remember.

ProtocolVariable mapGuideline selection
02

Sampling and power analysis

Frame definition, sampling strategy, inclusion and exclusion rules, and an a priori power calculation that tells you the sample size the claim needs — before recruitment, not after a reviewer asks. Delivered with the justification paragraph, written for your methods section.

Sampling planPower calculationG*Power output
03

Instrument development and validation

Questionnaires, scales and interview guides built from the construct up, then piloted and tested — content validity, reliability, and factor structure where the instrument is new or adapted. Adapting a published scale without revalidating it is one of the more common reasons a psychometric paper is returned.

Cronbach αCVI / CVRPilot report
04

Ethics and consent documentation

Ethics committee and IRB submissions, participant information sheets, consent forms, data management plans and the anonymisation approach — prepared in the format your committee expects, and consistent with the protocol rather than written separately from it.

IRB packConsent formsData management plan
05

Primary data collection

Fieldwork we run or supervise: online and offline surveys, clinical and institutional records, structured observation, in-depth interviews and focus groups, with enumerator training and live field monitoring. Response quality is watched while collection is open, when it can still be corrected.

Field teamResponse monitoringRaw dataset
06

Secondary data sourcing

Locating, acquiring and merging existing data: national statistical series, health and demographic surveys, World Bank and OECD indicators, Scopus and Web of Science exports, patent, market and institutional repositories — with licences and terms of use checked before anything is merged.

Source mapLicence checkMerged panel
07

Data cleaning and preparation

Coding, recoding, missing-data treatment with a stated mechanism, outlier handling, transformation and assumption testing — all recorded in a script rather than done by hand in a spreadsheet, so that every exclusion has a reason attached to it.

Cleaning scriptMissingness reportCodebook
08

Quantitative analysis

From descriptives to structural equation models, survival analysis, multilevel modelling, econometric panels and machine learning — chosen for the design and the data, not for what is fashionable in the field. Effect sizes and intervals are reported alongside p-values as a matter of course.

Analysis scriptOutput tablesFigures
09

Qualitative analysis

Transcription, coding frameworks, thematic and content analysis, grounded theory, framework analysis and mixed-methods integration, with an audit trail and intercoder agreement where the design calls for it — and a reflexive account where the guideline expects one.

Coding frameTheme mapAudit trail
10

Evidence synthesis and bibliometrics

Systematic search strategies, screening and extraction, risk-of-bias assessment, meta-analysis and meta-regression; plus bibliometric and scientometric mapping for reviews and research-performance work. The search strategy is written so a stranger can re-run it and get your numbers.

PRISMA flowForest plotsScience maps
11

Reporting, reproducibility and reviewer response

The methods and results sections written to the target journal's requirements, the reproducibility pack assembled, the repository deposit prepared, and statistical reviewer comments answered with the reanalysis they ask for rather than a paragraph explaining why it is unnecessary.

Methods textReproducibility packResponse letter

Environments

We work in your software, not only ours.

So that the files you receive open on your machine, and your student or co-author can run them again next year without asking us for a licence.

SPSSR · tidyverse, lavaan, metaforStata Python · pandas, statsmodels, scikit-learnSAS AMOSSmartPLSMplus JASPG*PowerNVivo MAXQDAVOSviewerBiblioshiny

Delivered as

  • Annotated syntax or script files
  • Cleaned dataset in an open format
  • Codebook and variable dictionary
  • Output tables in editable form
  • Figures at journal resolution

Turnaround

  • Scoping note: 1–2 working days
  • Methodology consultation: within a week
  • Single analysis: typically 2–3 weeks
  • Reviewer response: 5–10 working days
  • End-to-end study: quoted with the scope

Not on the list

  • We do not write introductions or discussions
  • We do not accept authorship
  • We do not fabricate or pad datasets
  • We do not guarantee a significant result
  • We do not guarantee acceptance by a journal

Not sure which of the eleven you need?

Describe the study in a paragraph. We will tell you which stage you are actually at — which is often earlier than people expect — and what it would take from there.

Send a brief See the nine stages

Methods library

Find the analysis your design actually supports.

Filter by what you are trying to establish, the kind of variable you are working with and the software you have. Each entry states when the method applies and what it produces. Use it to check our thinking, or to arrive at the first call already knowing what to ask for.

53Methods documented, across ten families
7Objectives, from comparing groups to mapping a field
14Software environments the entries are written for

The method follows the design. Not the other way round.

The most expensive mistake in applied statistics is picking a test because it is familiar, then discovering that the design cannot support it. The second most expensive is picking one after looking at the data, which turns an honest study into an accidental fishing expedition.

So the library is organised the way the decision is actually made: start from what you are trying to establish, narrow by the kind of variable you have, and only then look at what your software can run. If two methods survive that filter, the choice between them is a conversation — and usually a short one.

Nothing here is a substitute for a methodologist reading your protocol. It is a way to arrive at that conversation already knowing the vocabulary, and to check that what we propose is what your design implies.

Reading an entry

  • Family — the broad class the method belongs to: comparison, regression, causal inference, latent variable, qualitative, synthesis and so on
  • Objective — the question it answers: compare groups, test association, predict an outcome, estimate a causal effect, model latent constructs, interpret text, map a field
  • Data type — what the outcome variable looks like: continuous, categorical, ordinal, count, repeated measures, time-to-event, text, bibliographic records
  • When it applies — the design conditions under which the method is the right answer, and the assumption most often violated
  • Software — the environments we routinely run it in and can hand you working syntax for
0 methods match
Ask which one fits your study

The families, briefly

What each group of methods is for.

Ten families cover almost every design we see. If yours does not fit one of them, that is worth a conversation rather than a filter.

Comparison

Is the outcome different between these groups, or before and after? t-tests, ANOVA and their rank-based equivalents, chosen by how many groups there are and whether the same people appear twice.

Association

Do these two things move together? Correlations, chi-square, agreement statistics — descriptive claims that do not, on their own, license a causal sentence.

Regression

What predicts the outcome, holding other things constant? Linear, logistic, ordinal, count and multilevel models, with the choice driven by the shape of the outcome variable.

Causal inference

Would the outcome have differed had the exposure differed? Matching, differences-in-differences, instrumental variables, fixed effects and discontinuity designs — for observational data with a defensible identification strategy.

Time and survival

How long until something happens, and what changes it? Kaplan-Meier, Cox models and time-series forecasting, with censoring handled properly rather than dropped.

Latent variable

How do we measure something we cannot observe directly? Factor analysis, structural equation modelling, PLS-SEM, mediation and moderation — the backbone of most management and psychology submissions.

Multivariate

Too many correlated variables to look at one at a time. Component analysis, clustering and discriminant analysis, usually as a step towards a model rather than a finding in itself.

Qualitative

What do people actually say, and what does it mean? Thematic, framework, content and grounded theory analysis, with coding frames, audit trails and intercoder agreement where the guideline expects it.

Evidence synthesis and bibliometrics

What does the literature, taken together, support? Meta-analysis, meta-regression and risk-of-bias appraisal; plus co-citation and performance mapping for reviews and research-office reporting.

Environments we work in

Syntax you can run yourself.

SPSSR · tidyverse, lavaan, metaforStata Python · pandas, statsmodels, scikit-learnSAS AMOSSmartPLSMplus JASPG*PowerNVivo MAXQDAVOSviewerBiblioshiny

We work in your environment where you have one, so that the files you receive open on your machine and your student or co-author can run them again next year. Where you have no preference, we default to R for quantitative work and NVivo for qualitative, and supply the output in a form your co-authors can read without either.

Two methods survived your filter. Which one does your design imply?

Send the design and the variable list. You will get a short written answer naming the method, the assumption to check first, and what the reporting guideline expects you to state.

Ask a methodologist See the services

The method

Nine stages. Every one ends with something you can hold.

The sequence below is the whole of what we do. Nothing in it is unusual on its own — what matters is where each step sits, and which of them you have to sign before the next one starts.

9Stages, run in order, each with a named deliverable
3Gates you sign, placed before the expensive mistakes
3Checks before release, never by the same pair of hands

The order is the argument.

The analysis plan is Stage 3 and data collection is Stage 5. That single ordering is the difference between a study a statistical reviewer accepts and one they send back — because it means the tests were chosen while they were still a design decision, and not after somebody had seen which comparison looked promising.

It also removes the temptation that produces the problem, which is rarely dishonesty. It is deadline pressure meeting an ambiguous dataset at two in the morning.

Choosing a test after seeing the data is how p-hacking happens by accident.

What a stage has to produce

Every stage ends with something you can hold: a document, a script, a dataset or a certificate. A stage that produces only a conversation has not finished.

  • Stage 1 — a written scoping note, or a recommendation to proceed without us
  • Stage 3 — the analysis plan and power calculation, before collection
  • Stage 6 — the cleaned dataset, cleaning script and codebook
  • Stage 8 — a reproduction certificate from a second analyst
  • Stage 9 — methods text, reproducibility pack, deposit-ready dataset

Same activities, different order

Decide the analysis before you collect.

Pre-specification is free at the start and impossible to retrofit at revision. This is the entire reversal.

The usual order

  1. Design the study
  2. Collect the data
  3. Look at what you have
  4. Choose a test that fits it
  5. Discover the sample was too small

The sample size is justified after the fact, which a reviewer can see.

Our order

  1. Design the study
  2. Pre-specify the analysis
  3. Calculate the sample it needs
  4. Collect the data
  5. Run what the plan said

The plan is dated and signed, so the method is a record rather than a claim.

The sequence

Decide, gather, prove.

Three phases, nine stages. Scroll and the spine draws to whichever stage you are reading; the gates clear as their conditions are met.

DecideSTAGES 01 – 03

Everything expensive is settled here, before a participant is recruited or a dataset is licensed.

GatherSTAGES 04 – 06

The data arrives and is never touched by hand. Everything after it comes out of a script.

ProveSTAGES 07 – 09

Run it, prove it reproduces from your raw file, then write it up so a referee can follow it.

01

Scoping and feasibility

Decide

We read the proposal, the dataset description or the referee's comments verbatim, and tell you what the design can and cannot support. This stage is free, and it ends in writing whether or not you engage us — including, sometimes, with a recommendation to proceed without us.

You receiveWritten scoping noteScope and feeNamed analyst
02

Design and protocol

Decide

Variables, hypotheses, unit of analysis and comparison structure are fixed, and the reporting guideline the study will be written against is chosen now rather than at submission. The unit of analysis is written down even when it seems obvious — especially then, since most multilevel errors are a unit-of-analysis error nobody recorded.

You receiveProtocolVariable mapReporting checklist
03

Analysis plan and power

Decide

The primary analysis is named and justified against the design, the secondary and subgroup analyses are declared in advance, and the handling of missing data and outliers is decided by rule rather than by inspection. The sample size is calculated for the smallest effect worth finding.

Where the data already exists, we report the precision your sample supports instead of a post-hoc power calculation — which is mathematically redundant with the p-value and signals to a reviewer that the plan came after the result.

You receiveAnalysis planPower calculationJustification paragraph
Gate — nothing is fitted untilYou sign the analysis plan, and so does your supervisor or PI where one exists. After that, any change to the method is a numbered, dated amendment with a reason attached.
04

Instrument and ethics

Gather

Questionnaires, scales and interview guides built from the construct up or adapted, then piloted and tested for reliability and validity. A published scale translated or moved to a new population is a new instrument, and we revalidate it rather than cite the original.

Ethics and consent documents are prepared in your committee's format, and we check that the consent obtained actually covers the analysis and the deposit you intend — the clause nobody reads until a journal asks for the data.

You receiveValidated instrumentPilot reportEthics submission pack
05

Data collection or sourcing

Gather

Fieldwork we run or supervise, with enumerator training and response quality monitored while collection is open rather than audited after it closes — when nothing can be corrected. Or existing datasets located, licensed and merged, with provenance and access date recorded for every source.

Your raw file is copied once, set read-only and never opened again. Everything downstream reads from that copy.

You receiveRaw datasetField or sourcing logData query list
06

Cleaning and codebook

Gather

Coding, recoding, missing-data treatment with a stated mechanism, outlier decisions taken by the rule set in the plan, and assumption testing — all of it in a script rather than by hand in a spreadsheet. Every exclusion carries a reason, so when a reviewer asks why n fell from 412 to 388 there is a line to point at.

You receiveCleaned datasetCleaning scriptCodebookMissingness report
Gate — analysis does not start untilThe cleaning script runs from your raw file to the analysis dataset in one pass, in a clean session, with no manual step in between. A value typed directly into a table cannot be reproduced, and one is enough to make the Stage 8 certificate meaningless.
07

Analysis

Prove

The pre-specified analysis is run, with the assumptions your model actually requires tested and reported — not a standard battery that buries the check that mattered. Where an assumption fails we choose openly between transforming, switching to a robust alternative and respecifying, and record which and why.

Effect sizes and confidence intervals accompany every p-value, including for null results, and every specification fitted goes in the analysis log — not only the one that worked.

You receiveAnalysis scriptOutput tablesJournal-ready figuresAnalysis log
08

Independent re-run

Prove

A second analyst who was not on the project receives your raw file and the scripts — never the output tables, and never a summary of what to expect — and reproduces every reported number in a clean session. Anything that does not match is logged and resolved by fixing the script, never by adjusting the report to agree with it.

Reproduction runs from the raw file, not from the cleaned dataset. Starting from the cleaned data would test the analysis script alone and leave every cleaning decision unchecked.

You receiveReproduction certificateException logSigned guideline checklist
Gate — nothing is released untilEvery reported figure has reproduced exactly from the raw file, and the methods text has been checked against the reporting guideline for your design and signed by a third person who ran neither analysis.
09

Reporting and deposit

Prove

Methods and results written to the target journal's requirements in journal-ready past tense, with software and package versions named. Every number quoted in the text is cross-checked against the table it came from. The reproducibility pack is assembled, the repository deposit prepared with access conditions matching the consent, and the data availability statement drafted.

Delivered on a call rather than as a zip file, because most of the value transfers in that half hour — and because you cannot defend what nobody explained.

You receiveMethods and results textReproducibility packDeposit-ready dataset

Stage 8, in detail

The check most people do not run.

Ask anyone quoting for your analysis whether a second person re-runs it from your raw file before release. The answer separates a service from a supplier.

STAGE 08 Independent reproduction, from the raw file 0% match
1 · The analystRuns the pre-specified analysis and produces the script behind every table and figure.
2 · The reproducerReceives only your raw file and the scripts, and draws the same result independently in a clean session.
3 · The comparisonEvery reported figure is checked against theirs. A mismatch is fixed in the script, never in the report.

Illustrative. In practice 98% of analyses reproduce exactly on the first re-run; the remainder are returned to the client as a data query before anything is released. It is not an acceptance rate, and we do not present it as one.

The deliverable

What physically arrives at the end.

Not a slide summary of your results. A pack your co-author can open in five years and run again.

Cleaned datasetDocumented, version-controlled, in an open format
Analysis scriptsAnnotated, running end to end from your raw file
CodebookEvery variable, value label and derivation
Tables and figuresFormatted to your target journal, publication resolution
Methods and results textJournal-ready past tense, with software versions
Reproduction certificateSigned by the analyst who re-ran it, with the exception log
Reproducibility packAssembled for repository deposit, with session information
Data availability statementDrafted for your journal's requirement
Anticipated reviewer questionsThe three a referee will ask, answered in advance

Plus the acknowledgement wording, drafted for you. We are acknowledged; we do not take authorship, because statistical assistance does not meet the ICMJE criteria on its own.

Where people actually arrive

Most engagements do not start at Stage 1.

We take work at any point in the sequence. It costs less, and the result is stronger, the earlier it starts — but the honest position is that we are usually met in the middle.

ENTERS AT STAGE 01 A proposal or synopsis, no data yet Everything is still open. This is the cheapest engagement to run and the strongest to defend, because the plan genuinely precedes the data.
ENTERS AT STAGE 06 Data collected, analysis not begun Our most common entry point. Stages 2 and 3 are reconstructed and dated, and labelled in the report as a retrospective analysis plan — never presented as though they came first.
ENTERS AT STAGE 07 A reviewer has raised statistical comments Reanalysis against the referee's actual point, then Stages 8 and 9 in full, with a response letter that answers the comment rather than explaining why it is unnecessary.

A plan written after the data arrived is dated and labelled as such. It is the single easiest thing for a reviewer to catch, and the hardest to recover from.

Getting started

Four steps, no proposal deck.

We read it, and a methodologist replies

Within one business day, in writing, from the person who would actually run the work — not a sales team summarising it back to you.

NDA first, then the material

We send the NDA before you send anything. Only after it is signed do you send the protocol, the dataset description or the referee's comments. Never identifiable participant data by email.

A scoping call, free of charge

Thirty to forty-five minutes on what the design can and cannot support, ending in a written note whether or not you engage us.

A scope and a fee — or a recommendation to proceed alone

If a supervisor or a departmental statistician can answer the question, that is what the note will say. It happens often enough that we mention it here.

Every engagement starts at Stage 1, whatever stage you are at.

The scoping call is free, ends in writing, and occasionally ends with us telling you to save your money.

Book the call Read the integrity position

Who we work with

Two programmes. One method.

For the individual researcher with a study to finish, and for the institution that needs research capacity it does not have on staff. Both run the same nine stages, with every check placed before the point where fixing it becomes expensive.

FreeFirst scoping call, with a written note whether or not you engage us
1Named analyst on your project, from scoping to reviewer response
4Ways to engage, priced on the work rather than the deadline

The problem

The statistics reviewer is the one who stops you.

Language and structure are fixable in a week. A sample that cannot carry the claim, a test that assumes something your data violates, or an analysis nobody can reproduce are not fixable at revision — they send you back to collection, a year later, with the funding spent.

What we do about it

Take the decisions early, on the record.

You get a named analyst who reads the design before the data exists, an analysis plan you sign, and a second analyst who reproduces every number before it reaches you. When a referee raises a statistical comment, answering it is a week's work rather than a year's.

Where researchers usually start

Open a stage to see what happens in it and what you receive at the end.

Who this is for

  • PhD scholars and postdoctoral researchers
  • Faculty and principal investigators
  • Clinicians running audits, registries and trials
  • Thesis candidates whose committee wants the methods tightened
  • Authors with statistical comments to answer
  • Anyone preparing a dataset for a repository or a funder

How we engage

ModelSuitsIndicative
Methodology consultation
One session, written note after
Deciding design, sample size or testfrom USD 120
Single analysis
Fixed scope, fixed fee
One dataset, one paperfrom USD 450
End-to-end study
Design through deposit
Thesis chapters and funded projectsquoted
Reviewer response
Reanalysis and letter
Statistical comments from a journalfrom USD 300

Indicative only, in US dollars; billed in USD, EUR, GBP or INR. Priced on design complexity, dataset size and turnaround. The first scoping call is free and we will say if you do not need us.

What is included

  • A named analyst, not a queue
  • The nine stages from your entry point on
  • Independent reproduction before release
  • Journal-ready methods and results text
  • One revision round within 30 days

What is quoted separately

  • Reviewer response after delivery
  • A second paper from the same dataset
  • Re-analysis on a changed research question
  • Primary data collection we run in the field
  • Repository deposit beyond the prepared pack

What we will not do

  • Write your introduction or discussion
  • Accept authorship
  • Fabricate, simulate or pad a dataset
  • Guarantee a significant result
  • Guarantee acceptance by a journal

Turnaround

What to expect, and when the clock starts.

On complete inputs, not on the enquiry. Where a client is late supplying data or a decision, the delivery date moves by the same number of days — and we confirm that in writing on the day it happens, not at the end.

EngagementTargetClock starts
First reply to an enquiry1 business dayOn receipt
Scoping note after the call2 business daysEnd of the call
Methodology consultationWithin 1 weekSigned scope
Single analysis2–3 weeksComplete data and signed analysis plan
Reviewer response5–10 business daysReceipt of the referee report
End-to-end studyQuoted with the scopeSigned scope

Tell us where the study is, and we will tell you which programme fits.

Often the answer is the smallest one — a single consultation that changes the design before anything is collected.

Send your brief See the nine stages

Our promise

We change how the analysis is done. We never change what it found.

That line governs every project we take. We will tell you which test your design supports, run it correctly, and write the methods so a referee can follow them. What we will not do is move a threshold, drop a case that spoils a distribution, run the model until something reaches significance, or report the one specification of twelve that worked.

98%Analyses that reproduce exactly from the delivered script
3Separate checks, never performed by the same person
NDASigned before you send us anything at all

The position, in full

What you are actually buying when you buy an analysis.

If the analysis does not support the hypothesis, you get that result, with an honest reading of what it means and what a reviewer will make of it. A null finding you can defend is worth more than a significant one you cannot — and considerably more than one that gets retracted three years later.

This matters commercially, not only ethically. The market we sit in contains people who will produce a significant result on request. What they are selling is a liability with a p-value attached: the paper is accepted, then a reader cannot reproduce it, and the correction notice carries your name rather than theirs.

It is also why our name belongs in the acknowledgements rather than the author list. Statistical assistance and data curation do not meet the ICMJE authorship criteria on their own, and we will draft that disclosure for you — along with the data availability statement your journal now asks for.

Where an integrity question arises mid-project — an inconsistency in the raw file, a consent record that does not cover the analysis, a co-author asking for a case to be removed — we follow COPE guidance, tell you in writing, and stop work until it is resolved.

A significant result we cannot defend is not a deliverable. It is a future correction notice with your name on it.

Analyses that reproduce exactly from the delivered script0%

Every delivery is re-run from the raw file by a second analyst who was not on the project. If a number does not reproduce, it does not ship. The remainder are cases returned to the client for a data query before release.

Findings and effect sizesReported as computed
Model specifications triedAll disclosed, not just the best
Your raw dataNever altered, always returned
AuthorshipYours; we go in the acknowledgements

Separation of duties

Three checks, and the paperwork behind each one.

An analysis is only as trustworthy as the weakest check it survived. Ours are separated on purpose: the analyst who runs the model is not the analyst who reproduces it, and neither of them signs off the methods text.

Check one · the analyst

Runs the pre-specified analysis, tests and reports the assumptions, records every specification tried in the analysis log, and produces the script that generates each table and figure.

Check two · the reproducer

A second analyst who was not on the project takes the raw file and the script and re-runs everything independently. Any number that does not match is an exception, logged and resolved before release.

Check three · the methods reviewer

A third person, neither of the above, reads the methods and results text against the reporting guideline for the design — CONSORT, PRISMA, STROBE, COREQ or whichever applies — and signs it off.

Quality, ethics and confidentiality

The commitments in detail.

Written out because a supervisor, an ethics committee or a research office will sometimes ask to see them, and because a promise nobody can check is not worth making.

Analytical quality

  • Assumptions tested and reported, not assumed
  • Independent re-run of every analysis from the raw file
  • Methods text checked against the reporting guideline for the design
  • Effect sizes and confidence intervals reported alongside p-values
  • All specifications tried are disclosed in the analysis log
  • Intercoder agreement reported on qualitative coding

Research ethics

  • We do not fabricate, simulate or pad datasets under any circumstances
  • We do not ghost-write papers or accept authorship
  • No selective reporting, threshold shifting or outcome switching
  • Ethics approval confirmed before we touch human-subject data
  • Conflicts of interest and our role disclosed in your acknowledgements
  • COPE guidance followed if an integrity question arises mid-project

Confidentiality

  • NDA signed before you send anything
  • Identifiers stripped or pseudonymised on receipt
  • Encrypted storage and secure transfer; no email attachments
  • Access limited to the named project team
  • Your data never reused, resold or shown as a sample of our work
  • Documented deletion at the end of the agreed retention period

Reproducibility

  • Every result traceable to a line in a script
  • Version-controlled analysis files with a changelog
  • Codebook and variable dictionary supplied with the dataset
  • Reproducibility pack assembled for repository deposit
  • Data availability statement drafted for your journal
  • Files delivered in formats your co-authors can still open in five years

Requests we decline, and what we offer instead.

These come up often enough to be worth stating in advance. None of them is a judgement about the person asking; most arrive from someone under a deadline who has been told this is normal.

The requestWhat we do instead
“Can you make the result significant?”Report what the data shows, with the effect size and interval, and help you write a discussion that survives the null.
“Can you generate the data? The deadline moved.”Decline, without exception. Where a deadline is genuinely immovable, we help you scope a study that fits it honestly.
“Just drop the cases that break normality.”Test whether normality matters for your design at all, then choose a method that tolerates the data you actually have.
“Write the whole paper and put my name on it.”Write methods and results. Introduction, discussion and conclusion stay with you, because they are what authorship means.
“Add me as an author on your analysis.”Nothing to add — we take no authorship. We draft the acknowledgement instead.
“Can you guarantee acceptance?”No one can. We can make the methods section the part of your paper the reviewer does not argue with.

Send us something before you send us anything.

The NDA is signed first. Then the study, the dataset description, or the three sentences from a referee that started all this.

Start the conversation See who would run it

Our team

The people behind NextDataLab.

Statisticians, methodologists, qualitative researchers and data engineers, supported by an advisory board drawn from medicine, management, economics and education. For this audience, knowing there is a named biostatistician is most of the buying decision — so the names are here rather than behind a form.

4Operating leads, each owning one part of the sequence
4Advisory board members who review contested designs
1Named analyst attached to your project, start to finish

You deal with a person, not a queue.

Every project is assigned a named analyst at the scoping call, and that person stays with it through delivery and through whatever the reviewers send back. They are the one who answers your email, and the one whose reading of your design you can argue with.

Behind them sit two people you will hear from less: the analyst who independently reproduces the work from your raw file, and the methods reviewer who checks the write-up against the reporting guideline. Both are deliberately not the person who ran the analysis.

Where a design is contested — an unusual identification strategy, a psychometric instrument being adapted across languages, a trial with a protocol amendment — it goes to the advisory board member for that field before we commit to it.

Who touches your project

  • Named analyst — designs, runs and writes; your single point of contact
  • Reproducing analyst — re-runs everything from the raw file, independently
  • Methods reviewer — checks the text against CONSORT, PRISMA, STROBE, COREQ or whichever guideline applies
  • Discipline lead — signs off the design where the field has its own conventions
  • Advisory board — consulted on contested or unusual designs

Nobody outside that list sees your data. Access is scoped to the named project team and logged.

Operating team

Four leads, four parts of the sequence.

Each lead owns a stretch of the nine stages end to end, which is why a handover between them is a scheduled event rather than an email.

Name to be insertedHead of Methodology

Owns study design, sampling and analysis plans across all disciplines. Runs the scoping call on most projects, and is the person who will tell you the design cannot carry the claim.

Name to be insertedLead Biostatistician

Owns clinical and health-sciences analysis and reviewer response: trials, epidemiology, survival models, and the statistical comments that come back from medical journals.

Name to be insertedLead, Qualitative Research

Owns coding frameworks, mixed-methods integration and audit trails, and the intercoder agreement reporting that COREQ and SRQR reviewers look for.

Name to be insertedHead of Data Engineering

Owns sourcing, cleaning pipelines, version control and reproducibility — including the packs that go to repositories and the scripts your co-authors will run in five years.

Advisory board

Consulted on contested designs, and on the disciplinary conventions that decide whether a defensible analysis is also a publishable one.

Name to be insertedClinical research

Advises on trial design, ethics review and health data governance, including protocol amendments and registry requirements.

Name to be insertedEconometrics

Advises on causal identification, panel data and impact evaluation, and on when an identification strategy will not survive a referee.

Name to be insertedPsychometrics

Advises on scale development, cross-cultural validation and structural models, including measurement invariance across groups.

Name to be insertedResearch governance

Advises on data management policy, FAIR compliance and integrity, and on what accreditation reviews expect to see documented.

Replace the placeholder names and add photographs before the site goes live. Keep the role lines and the paragraph under each — for this audience, knowing there is a named biostatistician is most of the buying decision.

Discipline desks

Subject specialists behind the four leads.

The methods are general; the conventions are not. A defensible mediation model in management is written differently from a defensible one in clinical psychology, and reviewers notice.

Medicine and health sciencesClinical trials, epidemiology, public health, nursing, dentistry, pharmacy
Management and commerceOrganisational behaviour, marketing, finance, operations, entrepreneurship
Social sciencesSociology, psychology, political science, anthropology, social work
Economics and policyDevelopment, labour, health economics, impact evaluation, econometrics
EducationPedagogy, assessment, higher education studies, learning analytics
Engineering and technologyIndustrial engineering, computing, energy systems, quality and reliability
Agriculture and environmentAgronomy, food science, climate, sustainability, natural resources
Life sciencesBiotechnology, microbiology, genetics, ecology, veterinary science
Law and humanitiesEmpirical legal studies, media, linguistics, mixed-methods humanities work

What researchers say

Voices from the studies we have worked on.

Quote to be inserted.

Name, designationDepartment, university · Discipline

Quote to be inserted.

Name, designationDepartment, university · Discipline

Quote to be inserted.

Name, designationInstitution · Discipline

Carry across the names already used on NextPUBGlobal and paste each person's actual words, with permission to name them and their institution. For this audience an unattributed quote reads as invented.

Ask to speak to the analyst, not the account manager.

The scoping call is with the person who would run your study. If they are not the right fit for the design, they will say so and bring in whoever is.

Request a call How an engagement runs

Insights

Notes from the analysis desk.

Practical, source-checked writing on research data and statistics — the things nobody explains until a reviewer has already asked. Each note answers one question, states what the reporting guideline expects, and shows the two or three sentences you would actually put in the paper.

6Notes published, on the questions we are asked most
1Question per note, answered in full rather than surveyed
0Sign-up walls between you and any of them

Written for the moment before the mistake.

Almost every note here exists because the same question arrived three times in a month, usually from someone who had already collected their data. Power calculations after the fact. Normality tests treated as a gate rather than a diagnostic. Listwise deletion applied silently because the software default did it.

None of these are exotic errors. They are what a competent researcher does when nobody told them the convention has moved, and the cost of finding out from a referee instead of from a colleague is measured in months.

What a note contains

  • The question, stated as a researcher would actually ask it
  • What the current convention is, and when it changed
  • What the relevant reporting guideline requires you to state
  • The two or three sentences to put in your methods section
  • The mistake reviewers flag most often on this point

All notes

Six questions, answered properly.

Sample size

The power calculation your reviewer will ask for, and when it is too late to run it

Why a post-hoc power analysis convinces nobody, what an a priori calculation needs as input, and how to report it in three sentences.

8 min read · DesignRead the note
Assumptions

Your data violated normality. That is usually not the problem you think it is

When the assumption actually matters, when the central limit theorem covers you, and the three alternatives worth considering before you transform anything.

7 min read · AnalysisRead the note
Missing data

Listwise deletion is a decision, not a default

How to state a missingness mechanism, when multiple imputation is worth the complexity, and what to write in the methods either way.

9 min read · AnalysisRead the note
Mediation

Baron and Kenny is thirty years old, and reviewers have noticed

What replaced the causal steps approach, why bootstrapped indirect effects are now expected, and how to report a mediation model in 2026.

10 min read · Latent variableRead the note
Systematic review

Your search strategy has to be reproducible by a stranger

What PRISMA 2020 actually requires in the supplementary file, and the five omissions that get a review desk-rejected.

11 min read · SynthesisRead the note
Data sharing

Writing a data availability statement when you cannot share the data

Ethical and legal restrictions are an acceptable answer. Silence is not. How to write the statement that satisfies the journal and your ethics committee.

6 min read · GovernanceRead the note

Titles, reading times and summaries are placeholders written to show the section working. Replace with your published articles and point each card at its own page — the card markup already carries a link, a kicker and a metadata line.

What we write about

Five recurring subjects.

Roughly in the order they cause trouble.

Design and sampling

Sample size justification, sampling frames, inclusion rules, and the difference between a study that is underpowered and one that is simply small.

Assumptions and diagnostics

Which assumptions matter for which method, how to test them without turning the paper into a diagnostics report, and what to do when one fails.

Measurement

Adapting published instruments, reporting reliability properly, factor structure, and measurement invariance when you compare across groups or languages.

Reporting standards

What CONSORT, PRISMA, STROBE and COREQ actually require, where the checklists disagree with journal templates, and what belongs in the supplement.

Reproducibility and data governance

Codebooks, repository deposits, FAIR principles, data availability statements, and what to do when the data genuinely cannot be shared.

One note a month, when there is one worth sending.

No digest, no round-up, no newsletter about our newsletter. If a month passes without a question worth answering at length, nothing arrives.

Enter a valid email address.

Front-end demo — connect this to your mailing list provider before launch. Unsubscribe in one click, and we do not pass addresses to the publishing side of the business.

Address captured

Nothing has been sent — this form is not yet connected to a list. Wire it to your provider and this becomes the subscriber's confirmation.

A question that is not answered here?

Send it. If it comes up more than once, it becomes the next note — and you get the answer either way.

Ask the desk Browse the methods library

Start with a free methodology review

Send us the study, not a sales enquiry.

Share your proposal, your dataset description or the reviewer's comment. We come back with what the design supports, which analysis fits, what it would cost and how long it takes — or an honest note that you do not need us. We sign an NDA before you send anything.

1Business day to a reply, from a methodologist rather than a form
FreeFirst scoping call, with a written note after it
NDASigned before any material changes hands

What happens after you send this

Four steps, no proposal deck.

We read it, and a methodologist replies

Within one business day, in writing, from the person who would actually run the work — not a sales team summarising it back to you.

NDA, then the material

We send the NDA first. Only after it is signed do you send the protocol, the dataset description or the referee's comments. Never identifiable participant data by email.

A scoping call, free of charge

Thirty to forty-five minutes on what the design can and cannot support. It ends with a written scoping note, whether or not you engage us.

A scope and a fee, or a recommendation to proceed alone

If a supervisor or a departmental statistician can answer the question, that is what the note will say. It happens often enough that we mention it here.

Researcher enquiry

For scholars, faculty and clinicians.

Enter your name.
Enter a valid email address.
Choose the closest discipline.
Tell us where the study stands.
A few sentences is enough to start.

We reply within one business day. Do not attach identifiable participant data until the NDA is signed.

Enquiry captured

Nothing has been sent — this page is not yet connected to a mailbox. Wire the form to your backend or form service and this becomes the researcher's confirmation.

Institutional enquiry

For universities, departments and research centres.

Enter the institution name.
Enter a contact name.
Enter a valid email address.
Choose the closest option.

We will send an outline scope and a call invitation, not a brochure.

Request captured

Nothing has been sent — connect this form to your backend and this becomes the institution's confirmation, with its reference number.

Research enquiries

research@nextdatalab.com
Design, analysis and reviewer response.

Institutions

institutions@nextdatalab.com
Capacity, training and governance.

Speak to someone

+91 98186 45221
Weekdays, 10:00–19:00 IST.

Before you write

Questions we are asked before the first call.

No — and the earlier you come, the less it costs. A proposal, a research question and a rough idea of the population is enough for a scoping call. If anything, arriving with data already collected narrows what we can do for you.

Statistical and data support is a normal, disclosable part of research, and most institutions expect it to appear in the acknowledgements. We draft that disclosure for you. What is not allowed — and what we do not do — is ghost-writing, fabricated data or authorship in exchange for analysis. Tell your supervisor we exist; the response is usually relief.

The research question, the design, the approximate sample size, the software you use if you have a preference, and the deadline. If a referee has already commented, send the comments verbatim — they tell us more in three sentences than a summary does in a page.

NDA first, always. Identifiers are stripped or pseudonymised on receipt, transfer is encrypted, and we do not accept participant data as email attachments. Access is limited to the named project team, and everything is deleted at the end of the agreed retention period with the deletion documented. We confirm ethics approval before touching human-subject data.

No, and anyone who does is selling something else. What we can do is make the methods section the part of your paper a reviewer does not argue with, and make sure that if a statistical comment does come back, answering it is a week's work rather than a year's.

Usually the Dean of Research, the research cell, or whoever owns the accreditation file — but a single department can engage us on its own. Use the institutional form; we reply with an outline scope and a call invitation rather than a brochure.