Flag of Indiaसत्यमेव जयते
EconomyNCERT Class 11 · Statistics for Economics

Collection of Data

This chapter explains where statistical data come from (primary vs secondary sources) and how they are gathered (census vs sample survey), along with the instruments, modes, and sampling techniques that make the evidence trustworthy.

⏱ 7 min readGS-III7 sections5 memory tricks
Why this matters for UPSC

A perennial Prelims favourite for static-cum-current facts — Census 2011 numbers, the Registrar General of India, the Census Act 1948, and the census-versus-sample-survey distinction — plus the logic of random sampling and exit polls. For Mains it underpins GS-III economy themes on the reliability of official statistics, the NSSO/NSO statistical system, and evidence-based policymaking, and feeds GS-I population and demographic discussions. Mastering data-collection methodology helps you critically evaluate any data-driven argument in the exam.

Understand the chapter

Why Collect Data: Variables, Observations, Information

The whole purpose of collecting data is to provide evidence for a sound, clear solution to a problem. A characteristic that changes in value — like food-grain output across years — is a variable (denoted X, Y or Z); each value it takes is an observation, and the full set of observations is the data. Data is therefore a tool that converts raw variation into usable information.

  • Variable: a measurable characteristic that changes in value (e.g., year X, output Y).
  • Observation: a single value taken by a variable.
  • Data: a collection of observations yielding information about a problem.
  • Anchor figure: food-grain output rose from 108 MT (1970-71) to 272 MT (2016-17).

Sources of Data: Primary vs Secondary

Statistical data come from two sources. Primary data are collected first-hand by the investigator through an original enquiry; secondary data have already been collected, scrutinised and tabulated by some other agency. Crucially, the same data are primary to the source that first collects them and secondary to everyone who later reuses them, and secondary data save time and cost.

  • Primary data: first-hand, original enquiry by the researcher.
  • Secondary data: already collected and processed by another agency (government reports, books, websites).
  • Relativity: data are primary to the first collector, secondary to later users.
  • Trade-off: secondary data are cheaper and faster but may not fit the new study exactly.

The Survey and the Questionnaire

A survey is a method of gathering information from individuals, and its main instrument is the questionnaire or interview schedule, either self-administered by the respondent or put by a trained enumerator. A good questionnaire is short, easy to understand, ordered from general to specific, and free of ambiguity, double negatives, leading questions and suggested alternatives. Questions are either closed-ended (structured) or open-ended (unstructured).

  • Closed-ended: fixed options — two-way (yes/no) or multiple-choice; easy to score, harder to frame well.
  • Open-ended: free responses; richer but hard to interpret and score.
  • Any Other (please specify): captures answers the researcher did not anticipate.
  • Question hygiene: avoid double negatives, leading clues and offered alternatives that bias replies.

Modes of Data Collection and the Pilot Survey

There are three basic modes — personal (face-to-face) interviews, mailing questionnaire surveys, and telephone interviews — each with distinct cost, reach and bias trade-offs. Personal interviews give the highest response rate and allow all question types but are costliest; mailed surveys are least expensive, reach remote areas and preserve anonymity (best for sensitive questions) yet suffer low response; telephone interviews are cheap and quick but miss those without phones. Before the main survey, a pilot survey (pre-testing) on a small group exposes flaws in the questionnaire.

  • Personal interview: highest response rate, can watch reactions; most expensive and time-consuming.
  • Mailing: least expensive, only mode to reach remote areas, maintains anonymity; low response rate.
  • Telephone: cheap, quick, can clarify questions; fails where people lack telephones.
  • Pilot survey/pre-testing: checks question suitability, clarity, enumerator performance, cost and time.

Census vs Sample Survey

A census, or complete enumeration, studies every element of the population, whereas a sample survey studies a representative subset and infers about the whole. India's population census is decennial — a house-to-house enquiry covering all households, with demographic data published by the Registrar General of India, the last held in 2011. Most surveys are sample surveys because they deliver reasonably reliable information at lower cost and in less time, with smaller, better-trained and better-supervised enumerator teams.

  • Census: every unit covered; exhaustive but costly and slow.
  • Sample survey: representative subset; cheaper, faster, permits more intensive enquiry.
  • Population/Universe: the totality of items under study; must be identified first.
  • Representative sample: smaller than the population yet reasonably accurate.

Sampling Techniques: Random vs Non-random

Sampling is of two main types, random and non-random. In random sampling every unit of the sampling frame has an equal chance of selection — the lottery method, now computerised, or selection via random number tables — which removes selection bias. Exit polls illustrate random sampling, where voters leaving polling booths are sampled to predict results, and their occasional errors remind us that a sample yields an estimate, not certainty.

  • Random sampling: equal probability of selection; lottery method or random number tables.
  • Sampling frame: the complete list of all population units from which the sample is drawn.
  • Non-random sampling: units not selected by equal chance (investigator's discretion).
  • Exit poll: a live random sample of exiting voters; predictions can still go wrong.

India's Census and Statistical System (Static Anchors)

India's census draws its legal authority from the Census Act, 1948, and is conducted by the Office of the Registrar General and Census Commissioner of India under the Ministry of Home Affairs. Census is a Union subject — Entry 69 of the Union List in the Seventh Schedule. For sample-survey-based secondary data, the National Sample Survey Office, now the National Statistical Office under the Ministry of Statistics and Programme Implementation, is the key nodal body.

  • Census Act, 1948 is the parent law; Census is Union List Entry 69.
  • Conducted by the Registrar General and Census Commissioner of India (Ministry of Home Affairs).
  • First Indian census: 1872 (under Lord Mayo); first synchronous census: 1881 (under Lord Ripon).
  • NSSO (1950, founded on P.C. Mahalanobis's initiative), now NSO under MoSPI, supplies key sample-survey data.

Key terms

Variable
A characteristic that changes in value across observations, usually denoted X, Y or Z.
Observation
A single value taken by a variable.
Primary data
Data collected first-hand by the investigator through an original enquiry.
Secondary data
Data already collected and processed by another agency and then reused.
Questionnaire / Interview schedule
The survey instrument — a set of questions, self-administered or put by an enumerator.
Census (Complete Enumeration)
A survey that covers every element of the population.
Sample survey
Study of a representative subset of the population to infer about the whole.
Population / Universe
The totality of items under study to which the results are meant to apply.
Sampling frame
The complete list of all population units from which a sample is drawn.
Pilot survey
A pre-test of the questionnaire on a small group before the main survey.

Must-know facts exam-ready

  • The Census of India is decennial (every ten years); the last one was held in 2011.
  • Census 2011 population was 121.09 crore (102.87 crore in 2001; 23.83 crore in 1901).
  • India's population rose by more than 97 crore over the 110 years from 1901 to 2011.
  • Average annual population growth fell: 2.2% (1971-81), 1.97% (1991-2001), 1.64% (2001-2011).
  • Census demographic data are collected and published by the Registrar General of India.
  • Legal basis of the census is the Census Act, 1948; Census is Entry 69 of the Union List.
  • First Indian census was 1872 (Lord Mayo); first synchronous census was 1881 (Lord Ripon).
  • Three modes of data collection: personal interview, mailing questionnaire, telephone interview.
  • Two question types: closed-ended (two-way or multiple-choice) and open-ended.
  • Two sampling types: random and non-random; random gives every unit an equal chance (lottery method/random number tables).
  • Food-grain output: 108 MT (1970-71) to 272 MT (2016-17), with 252 MT in 2015-16.
  • NSSO (estd. 1950, P.C. Mahalanobis), now NSO under MoSPI, is a key sample-survey data source.

Timeline

  1. 1872First census of India conducted under Viceroy Lord Mayo (non-synchronous).
  2. 1881First synchronous, complete census of India under Lord Ripon; decennial series begins.
  3. 1901Census recorded India's population at 23.83 crore.
  4. 2001Census population 102.87 crore; decadal growth rate 1.97% (1991-2001).
  5. 2011Latest census; population 121.09 crore; decadal growth rate 1.64% (2001-2011).

Memory tricks remember it for good

Primary-Personal, Secondary-Someone-else
P-P = Primary data collected Personally/first-hand; S-S = Secondary data from Someone else who already processed it.
💡 Never confuse which data type is original versus reused.
PMT — Please Mail, Telephone
P = Personal interview, M = Mailing questionnaire, T = Telephone interview.
💡 Recall all three modes of data collection without missing one.
Mail is CRAS
C = least-Cost, R = Reaches remote areas (only mode), A = Anonymity maintained, S = best for Sensitive questions.
💡 Pin down why mailed questionnaires win despite low response rates.
Census = R-T-A
R = Registrar General of India publishes it; T = every Ten years (decennial); A = backed by the Census Act, 1948 (Union List Entry 69).
💡 Bundle the census's who, when and legal basis for Prelims.
Random sampling = ELF
E = Every unit has Equal chance; L = Lottery method (or random number tables); F = drawn from the sampling Frame.
💡 Define random sampling completely in one breath.

Traps to avoid

  • Primary vs secondary is relative, not absolute: the same dataset is primary to its first collector and secondary to every later user.
  • Census is NOT a sample survey: census covers every unit (complete enumeration); a sample studies only a representative subset.
  • Population (Universe) is the whole group under study, NOT the sample drawn from it.
  • Random sampling means equal chance of selection — it does NOT mean haphazard or arbitrary picking.
  • The census is decennial with the last in 2011 (not annual); the growth rate is falling even as absolute population rises.
  • Highest response rate belongs to the personal interview, while least expensive and best-for-sensitive is the mailed questionnaire — do not swap their merits.

Exam focus

🧠 Prelims angles

  • Census 2011 data points (121.09 crore; declining decadal growth of 1.64%) and the publishing body (Registrar General of India).
  • Static census facts: Census Act 1948, decennial nature, Union List Entry 69, first census 1872/1881.
  • Distinguishing Census (complete enumeration) from Sample survey, and Population from Sample.
  • Sampling types: random (equal chance/lottery/random number tables) vs non-random; exit polls as random sampling.
  • Primary vs secondary data classification and example-based identification.
  • Matching modes of collection (personal/mail/telephone) with their respective advantages and disadvantages.

✍️ Mains angles GS-III

  • Reliability of India's official statistics and the case for sample surveys over complete enumeration.Argue that sample surveys give faster, cheaper and more intensive data (NSSO/NSO), then balance against the census's exhaustiveness and unique uses.
  • Evidence-based policymaking: how the quality of data collection determines the quality of governance.Link primary/secondary sources, sound questionnaire design and pilot testing to credible, bias-free policy inputs.
  • The census as an instrument of governance and welfare targeting in India.Connect the decennial census (RGI, Census Act 1948) to demographic data that drives schemes, delimitation and resource devolution.
Practice Economy questions from this syllabus →

Last-minute revision tick as you recall

  • Data = observations of variables, collected to give evidence for decisions.
  • Primary = first-hand; Secondary = reused; the label is relative to the user.
  • Survey instrument = questionnaire/interview schedule; questions are closed-ended or open-ended.
  • Three modes: Personal (highest response), Mail (cheapest, remote, sensitive), Telephone (quick, needs a phone).
  • Always pilot/pre-test the questionnaire before the main survey.
  • Census = every unit, decennial, RGI, last 2011, Census Act 1948, Union List Entry 69.
  • Sample = representative subset; cheaper, faster, intensive; most surveys are sample surveys.
  • Random sampling = equal chance (lottery/random number tables); the exit poll is a live example.
  • Census 2011: 121.09 crore, with decadal growth falling to 1.64% (2001-11).

Distilled from NCERT Class 11 · Statistics for Economics for UPSC. Always cross-check facts with the original NCERT.