Most computer science final year project rejections at the topic-approval stage trace back to one problem: a machine learning or data-driven idea with no actual dataset behind it. Seven sources cover almost every Nigerian computer science project’s data needs, from an off-the-shelf Kaggle dataset to a Nigeria-specific NITDA or NCC release — what each one holds, and which project types it actually supports.
| Source | What it holds | Best for |
|---|---|---|
| Kaggle | Tens of thousands of ready-to-use datasets and competitions, most already cleaned | Supervised ML projects (classification, regression, image recognition) |
| UCI Machine Learning Repository | Classic, well-documented benchmark datasets used in academic ML research since the 1980s | Comparing your model against a known baseline |
| Nigerian government data: the National Bureau of Statistics (nigerianstat.gov.ng) and the Nigeria Open Data Portal (data.gov.ng) | Official statistical releases and government-published datasets across sectors — health, education, agriculture, finance | A Nigeria-specific analytics or dashboard project |
| NITDA / NCAIR | National ICT policy documents, digital economy reports, and periodic AI/data initiatives | Literature and policy grounding for an AI-in-Nigeria project, rarely a raw dataset |
| NCC (Nigerian Communications Commission) | Quarterly telecoms subscriber, tariff and broadband penetration statistics | A telecoms, network or mobile-app-adoption project |
| GitHub | Open-source datasets and code released alongside research papers | Reproducing or extending a specific published study’s method |
| Your own collected or scraped data | Data you generate yourself, by survey, sensor, or scraping a public website | A project with no existing dataset that fits your exact research question |
Kaggle: the default starting point for a supervised learning project
Kaggle hosts tens of thousands of public datasets, most already reasonably clean, alongside past and active competitions with published leaderboards you can benchmark your own model against. For a final year project on disease classification, credit scoring, sentiment analysis or image recognition, search Kaggle first — a dataset with 10,000+ rows and a clear target variable, already used in published competition solutions, gives you a credible baseline to compare your own model’s accuracy against in Chapter Four. The caution: a dataset with thousands of existing public notebooks analysing it invites a supervisor’s question about what your project adds beyond replication — state explicitly in your problem statement what your specific model, feature engineering choice, or evaluation angle contributes that existing public analyses do not.

UCI Machine Learning Repository: the benchmark standard
The UCI Machine Learning Repository, maintained by the University of California, Irvine, has supplied benchmark datasets to academic machine learning research since the 1980s — datasets like Iris, Adult Income and Breast Cancer Wisconsin appear in thousands of published papers, which makes UCI datasets useful specifically when your project’s contribution is a new algorithm or technique rather than a new problem domain. Because these datasets are so widely used, published accuracy figures exist for dozens of algorithms against them, giving you a direct comparison point: “our model achieved X% accuracy against a published baseline of Y%” is a stronger Chapter Four sentence than an unanchored accuracy figure with nothing to compare it to.
Nigerian government data: NBS and the open data portal
The National Bureau of Statistics (nigerianstat.gov.ng) publishes Nigeria’s official statistical releases, and the Nigeria Open Data Portal (data.gov.ng) was set up to publish government datasets across health, education, agriculture, finance and other sectors. The portal is not always reachable — it could not be opened when this guide was checked in September 2026 — so treat NBS and the publishing ministry or agency’s own site as your first stop, and cite the agency that actually released the data. Coverage and update frequency vary considerably by dataset and by publishing agency — some datasets are current, others have not been refreshed in several reporting cycles — so check the dataset’s own publication date before building a project’s data-collection chapter around it, and state that date explicitly in your methodology rather than implying the data is current when it may not be. A dashboard or predictive-analytics project scoped around one specific, recently updated Nigerian government dataset reads as more grounded to a Nigerian panel than an imported dataset from a country your project never otherwise engages with.
NITDA and NCAIR: policy grounding, not usually a raw dataset
The National Information Technology Development Agency (NITDA) and its National Centre for Artificial Intelligence and Robotics (NCAIR) publish Nigeria’s digital economy and AI-related policy documents, including the National Artificial Intelligence Strategy — useful for grounding your background and literature review in the actual national policy context your project sits inside, and for justifying why an AI or data-driven project matters specifically to Nigeria’s stated digital economy goals. These sources rarely hand you a structured, ready-to-train dataset the way Kaggle or UCI do; treat them as your Chapter One and Two policy citations, not your Chapter Three data source.

NCC: the source for a telecoms or connectivity project
The Nigerian Communications Commission publishes quarterly industry statistics — subscriber numbers by operator, teledensity, broadband penetration and tariff data — the standard source for any final year project analysing network adoption, telecoms market structure, or connectivity as a variable in a wider study. Our guide to what the NCC’s own data actually shows about student internet access covers exactly how to read these figures without over-claiming what they measure — NCC subscriber counts, in particular, measure active SIM registrations, not unique individuals, a distinction that matters if your project’s variable is “number of internet users” rather than “number of active subscriptions.”
GitHub: reproducing or extending a published method
When your project’s contribution is extending, comparing, or re-implementing a specific published paper’s method — a common and defensible final year project shape — the paper’s own GitHub repository, if the authors released one, gives you the exact dataset split, preprocessing code and evaluation metric they used, which makes your comparison legitimate rather than an apples-to-oranges claim. Search the paper’s title plus “GitHub” or check the paper’s own “code availability” section; not every paper releases one, and building a project around reproduction only works if the original code and data are actually accessible.
Collecting or scraping your own data
When no existing dataset fits your exact research question — a Nigeria-specific sentiment analysis on a local social media discussion, a custom sensor dataset for an IoT project — collecting your own data is legitimate and sometimes the only real option. Web scraping public data is common for final year NLP and sentiment-analysis projects, but two things need stating plainly in Chapter Three: the site’s terms of service on scraping (many explicitly prohibit it, and a project should not depend on violating a platform’s own terms), and the Nigeria Data Protection Act 2023 if any scraped content includes personal data about identifiable individuals — anonymise or aggregate before storing and analysing. A sensor or survey-based dataset you collect yourself needs the same population, sample-size and instrument documentation as any other final year project’s primary data.
How does your data source connect to your limitations section?
Whichever source you choose shapes your limitations section directly, and naming that connection explicitly is worth doing before your panel asks. A Kaggle or UCI dataset limits your findings to whatever population and time period the original collectors sampled — your model’s performance says nothing about how it would generalise to a different context unless you state that boundary plainly. A self-collected or scraped dataset limits your findings to whatever sample size and collection window you achieved, which is usually smaller and narrower than a large public dataset. Our guide to what goes in the limitations section of a computer science project covers this in full — read it alongside your data-source choice, because the two sections should echo each other rather than read as unrelated afterthoughts.
How does your data source choice interact with your system development methodology?
If your project is a software-development or system-build project rather than a pure data-analysis one, your data source still matters — a system that classifies, recommends or predicts needs training data before it can be evaluated, and the methodology chapter should name where that data came from with the same rigour as a pure analytics project. Our comparison of system development methodologies for a computer science project covers how SSADM, Agile and prototyping structure the build process itself; the data source feeding your system’s intelligent component sits inside whichever methodology chapter you choose, documented with the same source, collection date and preprocessing detail as a standalone data-analysis project.
How do you cite a dataset in your reference list?
Treat a dataset the same way you would treat any other source: name the creator or publishing organisation, the dataset’s title, the year of publication or last update, and the URL or DOI where you accessed it — a Kaggle dataset citation names the uploader and dataset title, a UCI dataset citation names the repository entry, and a government dataset citation names the publishing agency (NCC, NBS or the relevant ministry) directly. Our general guide to referencing your project the way your Nigerian university actually requires covers how your specific department’s referencing convention handles a dataset citation, which sits closer to a “software or dataset” reference type than a standard journal article in most style guides.
How much data is enough for a machine learning project?
There is no single number — it depends on your model’s complexity and the number of features, not a fixed row count — but a dataset with fewer than a few hundred rows is a genuine constraint most Nigerian undergraduate departments will ask about at defence, because a small dataset makes overfitting hard to rule out. State your dataset’s size explicitly in Chapter Three, alongside your train-test split ratio and whether you used cross-validation, which matters more for a small dataset than for a large one. A dataset in the low thousands of rows, with a clear target variable and reasonable class balance, is a comfortable working size for most undergraduate classification or regression projects.
Frequently asked questions
Can I combine data from two different sources for one project?
Yes, and this is common — combining an NCC connectivity dataset with a survey you collect yourself, for example — but document each source separately in Chapter Three, with its own citation, collection date and any preprocessing you applied before merging.
Is it acceptable to use a synthetic or generated dataset?
For some project types (testing an algorithm’s behaviour under controlled conditions, for instance), yes, but state clearly that the data is synthetic and explain why — a supervisor expecting real-world data will ask why a synthetic dataset was chosen if this is not stated upfront.
Does Kaggle count as a credible academic source for my references?
Cite the dataset’s own documentation and, where one exists, the paper or report that originally published it, rather than citing “Kaggle” alone as the source — this is the same standard applied to any secondary dataset.
What if my chosen dataset turns out to have data quality problems?
Document the cleaning steps you took (missing value handling, outlier treatment, deduplication) explicitly in Chapter Three — a dataset that needed real cleaning work, honestly reported, is a stronger methodology section than one that implies the data arrived perfectly ready to use.
Can I scrape data from a Nigerian government website if it has no public API?
Check the site’s terms of use first; many government sites do not explicitly prohibit reasonable, non-disruptive scraping of public information, but confirm rather than assume, and avoid any scraping that could be read as placing excessive load on a public service’s servers.
How recent does my dataset need to be for a 2026 project?
There is no fixed cutoff, but state your dataset’s actual publication or collection date plainly, and if it predates recent, relevant changes in your project’s domain, address that limitation explicitly in Chapter Five rather than leaving it unstated.
Tesify helps you write the data-sources justification in Chapter Three clearly, cite each source correctly, and document your cleaning and preprocessing steps the way a panel expects to see them.
