Most of what gets written about I/O work skips straight to the finding. Somebody ran an analysis, something turned up, and now there's a recommendation. The part before that gets almost no attention, and it's where the majority of the hours go and where most of the mistakes are made.
Something has changed about that part, and it's worth naming clearly. Not long ago the hard bit was execution. Could you remember which test to run, could you get the syntax right, could you make the join work without breaking something. That barrier has dropped through the floor. A person with no coding background can now get working analysis code in a few minutes.
What hasn't changed is knowing which question to ask, knowing whether your data can honestly answer it, and knowing whether the answer that comes back is sensible. The scarce skill moved from operating the tools to judging the output. That's a good trade for people trained the way we were, because judgment about behavioral data is exactly what a psychology program builds and exactly what the tools can't supply.
A note on whether this applies to you. Depending on the role, you might run analyses constantly or almost never. It belongs in your toolkit either way, because even if you never open a dataset you will read analyses, commission them, and at some point sit across from a vendor quoting a number at you. All three require the same judgment as running one yourself, and there's no way to build that judgment without having done the work at least a few times.
What follows is how to work out what your data can answer, how to get it and get it usable, what the tools are actually for, and how to use AI on all of it without embarrassing yourself.
Part 1: What your data can actually answer
Whatever data exists was gathered for another purpose, by people who weren't thinking about the question you've been handed. So there are always two questions in play: the one you were asked, and the one your data can actually support. The gap between them is where most bad analysis lives.
This step gets skipped constantly and the reason is understandable. Pulling data feels like progress, and asking questions about the request feels like delay. The delay is much cheaper than the alternative.
Four questions to ask whoever asked you
Requests arrive pre-shaped. Nobody asks you to help them understand attrition. They ask you to pull the turnover numbers, and that pull is already somebody's solution to a problem they haven't described. These four walk it back, and they take one conversation.
- What decision does this feed, and when is it made? A readout that lands after the decision is a document. This also tells you how much precision is worth buying.
- What would you do differently depending on what I find? The most useful of the four, and the one people find hardest to answer on the spot.
- What do you already think is going on? You need their hypothesis so you can test it rather than accidentally confirm it, and so you know how far your answer sits from where they're starting.
- Has anyone looked at this before? There is often a half-finished analysis from eighteen months ago, and whoever ran it knows where the data is broken.
You won't always get all four answered and one of them is enough to change the work. The second question saves the most time. If every possible finding leads to the same action, there's nothing here to analyze, and what's wanted is usually reassurance or ammunition for a decision already made. Both are legitimate requests, and knowing which one you're serving changes what you build.
Common asks, and what they actually need
The requests are fairly repeatable across roles and organizations. For each one below: what you'd need, what it can answer, and where it stops. The last line is the one that matters, because that's where honest scoping happens and where you'll get caught out if you skipped it.
"Why are people leaving?"
Need Hire and termination dates, department, level, manager ID, voluntary and involuntary flag, internal transfers flagged separately.
Answers Where and when turnover concentrates, whether it has moved, which groups are most affected.
Stops at Why. Exit reason codes are selected by whoever ran the conversation, and "personal reasons" dominates everywhere. Getting to why requires exit interview text or a stay survey, which is a separate collection effort.
"Is our hiring process working?"
Need Applicant records with source, stage progression, assessment and interview scores, offer and hire decisions, plus some outcome for the people you hired.
Answers Where candidates drop out, how long stages take, whether your assessments agree with each other, and whether scores relate to later performance if outcome data exists.
Stops at How the people you rejected would have performed. Nobody ever finds out, which means any validity figure you calculate in-house is built only on the people who made it through and understates the real relationship. Worth saying out loud when you report one.
"Are our managers any good?"
Need Manager IDs linked to their teams, plus team-level outcomes such as turnover, engagement, or performance distribution.
Answers Whether outcomes vary between managers more than chance would predict, and who sits at the extremes.
Stops at Causation. Managers are not randomly assigned to teams. Bad numbers can mean an inherited mess, a failing product line, or a team nobody else wanted.
"Is engagement getting worse?"
Need Item-level responses across at least two waves, identical item wording, and response rates broken out by group.
Answers Change over time on matched items, and where that change is concentrated.
Stops at Any comparison at all if the items or the scale changed between waves. Matched wording is the entire basis of a trend, and items get quietly revised more often than people admit. It also tells you nothing about the people who didn't respond, and the reason someone skips a survey is rarely random.
"What are people actually saying in the comments?"
Need The raw open-ended responses rather than the platform's summary, the exact question that prompted them, and how many people answered out of how many were asked.
Answers What people chose to raise, which concerns recur, and the language they use for them — which is frequently more useful than the counts.
Stops at Prevalence. Only some people wrote something and they had a reason for doing it, so you know what was raised rather than how widely it's felt. Reporting theme counts as if they were percentages of the workforce is the most common mistake in this kind of work.
"Did the training work?"
Need Attendance records with dates, a behavioral or business outcome measured before and after, and ideally a comparison group who didn't attend.
Answers Whether the outcome moved for attendees relative to others.
Stops at Attribution, without a comparison group. Satisfaction scores answer a different question and are frequently offered as if they answer this one.
"Are we paying people fairly?"
Need Compensation, grade or level, location, tenure, performance rating, and demographic fields if you are permitted to hold them.
Answers Whether pay differences persist after accounting for legitimate factors.
Stops at Explaining any gap that remains. A residual difference tells you something is unaccounted for without telling you what it is. This one also carries real legal exposure — findings can be discoverable — and it should not start without counsel involved.
The pattern across all seven is the same, which is what makes it useful on questions that aren't on the list. Organizational data is reliably good at telling you where and how much. It runs out at why, and it runs out at whether this caused that. Anything you're asked that sits on the far side of those two lines needs either a separate collection effort or a plain statement that you can describe the pattern without explaining it. Saying that out loud early is far more comfortable than being asked for the causal claim in a readout and improvising one.
Decide what the results would mean before you look
Write down, in advance, what you would recommend under each plausible outcome. If turnover is concentrated in two teams, what then? If it's evenly spread, what then? If the effect is real but small, what then?
Two minutes of this does more for the quality of an analysis than any statistical choice you'll make later. If you can't fill in the branches, you aren't ready to look, and you'll be vulnerable to finding whatever you already suspected.
Part 2: Getting the data out of the systems
Getting data is a relationship problem before it's a technical one. The people holding it are busy, they field vague requests constantly, and they're usually accountable for that data under a policy somebody else wrote. Making their job easy is most of the trick.
Where it lives, and who owns it
- HRIS — Workday, SAP SuccessFactors, BambooHR, ADP, Dayforce. Headcount, job, level, location, manager, hire and termination dates. Usually held by HR Operations or People Analytics.
- Applicant tracking — Greenhouse, Lever, iCIMS, SmartRecruiters, Workday Recruiting. Applicants, sources, stages, interview records, offers. Held by Talent Acquisition.
- Survey platform — Qualtrics, Culture Amp, Microsoft Viva Glint, SurveyMonkey, Medallia. Engagement and pulse responses. Most of these have a reporting layer that shows you summaries and makes the underlying responses harder to reach.
- Performance and talent — Sometimes inside the HRIS, sometimes separate. Ratings, calibration outcomes, goals, succession lists.
- Learning system — Cornerstone, Docebo, Workday Learning. Reliable on who attended what and weak on almost everything else.
- Payroll and compensation — Salaries, bonuses, grade structures. Held tightly and usually behind a different approval than everything else.
- HR case management — ServiceNow, Zendesk, or whatever the service desk runs on. Employee relations cases and query volumes. Consistently underused and often the most revealing thing in the building.
Most real questions need data from more than one of these. Finding out whether your assessments predict anything means pulling candidate scores from the applicant system and performance outcomes from somewhere else, then matching them up person by person. That matching step is where the trouble usually starts, because two systems rarely share an identifier and the same person can appear differently in each one.
What you won't be given
Survey platforms typically enforce a minimum group size before they'll show a breakdown, usually somewhere between five and ten responses, so individuals can't be identified from a cut of the data. Compensation and case data often sit behind a separate approval, and sometimes behind a legal one.
These limits are worth respecting rather than working around. Pushing on a survey threshold is a fast way to lose the trust that makes the next survey worth running.
External data worth knowing about
Internal numbers mean very little without something to compare them against.
- O*NET OnLine — Free data on more than 900 occupations: tasks, skills, abilities, knowledge, work activities, work context. The full database is downloadable from the O*NET Resource Center under a Creative Commons licence. The obvious starting point for job analysis and competency work, and it saves weeks.
- BLS JOLTS — Monthly US quits, hires and separations rates by industry. What you reach for when somebody asks whether your turnover is bad, since "bad" only means something against a comparison.
- BLS Occupational Employment and Wage Statistics — Wage data by occupation and geography.
- Your own history — Two years of your own numbers is often the most relevant comparison available to you.
Both BLS sources are US-only. Most countries have a statistics office publishing equivalents.
One caution on benchmarks. A comparison built from organizations nothing like yours is worse than no comparison at all, because a number presented as a benchmark acquires an authority it hasn't earned and becomes very hard to argue with later. Check what's actually in the sample before you quote it.
Asking for it
Vague requests produce long email threads. Something closer to this gets you an extract:
Hi [name], I'm looking at [question] for [requester], with a readout on [date].
Could I get an extract covering [population] between [start date] and [end date], with these fields: [list them].
Two definitions so we're consistent: I'm treating [term] as [definition], and I'd like internal transfers flagged separately from exits.
CSV is ideal. If any field is restricted, tell me what's available instead and I'll work around it.
Ask for one level rawer than you think you need. A summary table answers exactly one question and blocks every follow-up, and the follow-up is usually the interesting part.
What comes back
What arrives depends on who you asked and what they assumed you needed. Roughly ordered by how much work each one creates:
- CSV or a clean tabular export — What you want. Ask for it by name.
- Excel with merged cells, stacked header rows, or color used to mean something — Needs manual repair before anything else can happen, and the meaning of the colors lives in somebody's head.
- A pre-aggregated system report — The numbers are already summarized, so you can't re-cut them. Go back and ask for the underlying rows.
- A dashboard with no export button — There is nearly always an extract behind it. Ask the system owner rather than screenshotting.
- PDF — Push back rather than retyping. Somebody generated it from something.
- Direct database or API access — Rare, and worth learning a small amount of SQL for if you ever get it.
The recognizable messes
Cleaning takes most of the time on a project and receives none of the credit. The problems repeat, which is the good news, because you stop being surprised by them somewhere around the third dataset.
- The same person under several IDs — Name changes, relocations, rehires, system migrations. Reconcile on a stable key and check for rehires before you assume a duplicate is an error.
- Two systems that disagree — Pick one as the source of truth, write down which and why, and stop reconciling case by case.
- Forty job titles covering six jobs — Group to a level or family before analyzing anything.
- Dates that mean different things — Confirm the definition of every date field. "Start date" can be the offer, the contract, or the first day at a desk.
- Free text where a category belongs — Standardize before counting. This is a good task to hand to a model.
- Cells that collapse when you segment — 1,200 employees looks like plenty until you cut by department and level and tenure and find yourself looking at four people. Decide a minimum group size before you start cutting and hold to it.
- Missing data that isn't missing at random — Check whether blanks correlate with anything. Empty performance ratings usually mean new joiners and leavers, who are frequently the exact people your question is about.
Every item on that list is a judgment call rather than a technical step. Someone competent working from the same file could reasonably decide differently on most of them, and the choices compound. That's what makes a decision log worth the trouble.
Keep a decision log
Every judgment call you make while cleaning is one you'll be asked to defend later, usually by the person whose team the finding concerns. A four-column sheet is enough: date, what you decided, why, and what it affects. It takes minutes while you work and saves the analysis afterwards.
Pay attention while you're in here, too. The finding often shows up during cleaning rather than after it, which is when you notice that most of the duplicate records come from one location.
Part 3: The tools, honestly
There's no single right answer here, and there's a lot of noise from people insisting their tool is the serious one. What follows is what each is actually for.
- Spreadsheets (Excel, Google Sheets) — Where most I/O work genuinely happens, and there's no shame in that. Good for looking at data, pivot tables, quick cuts, and anything you need to hand to a person who'll open it themselves. They struggle when you're repeating the same steps every month, when files get large, and when someone asks how you got a number. The failure mode is a workbook nobody can reproduce, including you in six months.
- SPSS — What most of us learned on. Menu-driven, strong for standard inferential tests and psychometrics, and it produces output that looks like what a journal expects. It costs real money and plenty of employers outside universities and large consultancies don't have a licence. Worth knowing that what you learned in SPSS transfers as concepts even when the software doesn't follow you into a job.
- R — Free. Strongest of these for statistics and psychometrics specifically, with packages built for exactly the things this field does. The real advantage is reproducibility: a script re-runs and gives the same answer, which matters when someone asks you to redo the analysis with one group excluded.
- Python — Free. Better than R for pulling data out of systems, handling text at scale, and automating anything repetitive. For straight statistical analysis the two overlap heavily and either is fine.
- SQL — Worth separating from the others, because it isn't analysis. It's retrieval, the language for getting data out of a database. A small amount goes a long way, and for many practitioners it's the highest-return thing on this list, because it removes your dependency on somebody else running an export for you.
Why people bounce off the coding ones
Plenty of I/O people try R or Python, get partway through a course, and quietly stop. The usual reason is that they set out to learn the language rather than to solve one specific problem with it. Courses teach syntax in an order that makes sense for programmers, and by the time anything useful appears you've lost interest.
The approach that works better is to take one thing you already do in a spreadsheet, do that exact thing in R, and let everything else arrive as you need it. This is also where AI has changed the picture most, because you can now get working code before you fully understand it, then read it line by line and learn from something that's already producing an answer you care about.
Free places to learn
- Kaggle Learn — Short hands-on courses in Python, SQL, and visualization. Free, and you can finish one in an afternoon.
- freeCodeCamp — Including a full Data Analysis with Python course with a certificate. Free.
- R for Data Science, 2nd edition — By Wickham and Çetinkaya-Rundel. The standard R text, free to read online.
- StatQuest — For statistical intuition. Worth it even if you already know the tests, because it explains what they mean rather than how to run them.
- Google's Data Analytics Professional Certificate — On Coursera. Free to audit all the videos and readings, around $49 a month if you want the certificate itself, with financial aid available.
- Google Colab — For running Python without installing anything on your machine.
Part 4: Using AI on all of this
This is where the barrier dropped, and it's worth being deliberate about rather than just pasting things in and hoping.
Confidentiality comes first
Employee data is personal data. Names, identifiers, salaries, performance ratings, exit reasons, and especially free-text comments, where people describe their manager in ways they'd never say out loud.
Use whatever your organization has actually approved, with an agreement in place, rather than a personal account. Check what's permitted before you paste anything, because "I didn't know" is not a position you want to be in when it comes up.
The practical technique that solves most of this: describe the structure instead of sharing the contents. You can get the great majority of the value by telling a model your column names, your row count, and what you're trying to learn, without any real records leaving your machine. When you do need to show it data, strip identifiers and replace them with sequential IDs, or build a synthetic sample of ten fake rows with the same shape.
Free text carries the most risk, because people write identifying details into comment boxes without noticing they've done it.
A simple test: if you wouldn't email the file to someone outside your organization, don't paste it into a model outside your organization.
What a good prompt actually contains
The point is the shape rather than the specific words, so here are the parts that make the difference.
- The structure of your data — Column names, what type each one is, roughly how many rows, and one row of invented sample values. Skip this and the model guesses at your structure and hands you code that fails immediately.
- What you're trying to learn, in plain language — Resist naming the test you think you need. Say what you want to know and let it propose an approach, then check that proposal against your own training. You'll catch bad suggestions, and occasionally you'll get a better one than you had in mind.
- Where you're working — Excel with no add-ins, Google Sheets, SPSS, R in RStudio, a Colab notebook. The answer is completely different for each, and a generic response gets you something that doesn't paste anywhere.
- What you already know is wrong with the data — Small groups, missing values, two systems that disagree, dates in three formats. Anything you'd mention to a colleague.
- A request for its assumptions — Ask it to state what the approach assumes and what would make the result misleading. This is the part almost everyone skips and it's the part that catches errors.
Three worked examples
Quantitative — turnover
I have a spreadsheet of 1,240 employees with these columns: employee_id (text), department (text, 9 values), hire_date (date), termination_date (date, blank if still employed), job_level (text, 5 values), manager_id (text). I want to know whether first-year turnover is meaningfully higher in some departments than others, and whether any difference is large enough to act on. I'm working in Excel with no add-ins. Three departments have fewer than 30 people. Tell me what approach you'd take, what it assumes, and what would make the result misleading. Then give me the steps.
Quantitative — survey data
I have engagement survey results from 480 respondents across 6 functions: 12 items on a 1 to 5 scale, plus tenure band and job level. Two functions have fewer than 15 respondents. I want to know where to focus and whether the differences between functions are real or noise. I'm working in SPSS. Walk me through what you'd do, and flag clearly where the small groups create something I should not be reporting at all.
Qualitative — open-ended comments
I have 412 free-text responses to the question "What one thing would most improve your experience here?" I want a coding scheme that covers them and gets applied consistently, with counts per theme. Start by proposing a set of themes based on a sample of 40 responses, tell me where categories overlap or are ambiguous, and wait for me to approve the scheme before you code the rest. Keep quotes verbatim so I can check your coding against the source.
That third one is worth dwelling on, because it mirrors how qualitative analysis is supposed to work anyway. A scheme gets proposed, a human checks it, then it gets applied consistently. Doing it in that order keeps you in the position of making the judgment calls, which is where you should be.
Checking the work
The failure mode is a confident, well-formatted, wrong answer. A few things that catch it.
Run it first on a small subset where you already know the answer. If it can't reproduce something you can verify by hand, don't trust it on the part you can't.
Ask for the numbers behind any claim, including the group sizes. A difference between two groups means something different when one of them has six people in it.
If it produced code or formulas, read the filters. Most errors live in what got included or excluded rather than in the statistics.
And ask it to argue against its own conclusion. "What would have to be true for this result to be an artifact rather than a real effect" is a question that surfaces problems remarkably often.
Where this leads
None of this is the interesting part of the work, and all of it determines whether the interesting part is any good. A finding built on a definition nobody agreed, or a cut of data too small to support it, falls apart the first time somebody looks closely, usually in public.
The barrier to doing this well is genuinely lower than it was five years ago, which means the excuse for avoiding it is thinner. You don't need to become a data scientist, and you do need enough fluency to know what your data can answer, enough patience to get it into shape, and enough judgment to tell when a confident answer is wrong.
Then the harder question arrives, which is what to do with what you found and how to get anyone to act on it. That's a separate skill and it's covered in Telling the Story.