A strong data science or AI final year project topic names a specific technique, a specific Nigerian dataset or problem, and a measurable output — for example, a fraud-detection model tested on anonymised transaction records, or a crop-yield predictor built on rainfall and soil data for one state. Below are topics grouped by eight sub-areas, each with a one-line research question, plus where the data for each one actually comes from.
What Makes a Data Science or AI Topic Strong Enough to Defend?
In some Nigerian computing departments, Data Science and Artificial Intelligence are offered as their own programme or specialisation track, separate from a general software-engineering final year project; in others they are simply a project theme. Either way, a topic in this space is defensible when it answers three questions before you write a single line of code: which named technique or model family will you use (not “AI” as a vague label, but a stated algorithm or model type), which dataset will you train and test it on and where that data actually comes from, and what single, measurable output will your Chapter Four report (an accuracy/F1 score, a forecast error, a classification result) rather than a vague claim that the system “works well.”
A useful test before you commit: can you write, right now, one sentence that names the algorithm, the dataset and the metric together — for example, “a random forest classifier trained on 400 anonymised loan records, evaluated by F1 score against a logistic-regression baseline”? If you cannot fill in all three blanks yet, the topic still needs narrowing, no matter how exciting the general idea sounds. This single-sentence test also doubles as the first line of your research proposal, so the work is not wasted.
What Data Science and AI Project Topics Are Available, by Sub-Area?
Thirty-two topics across eight sub-areas, each phrased as a researchable question. Pick one you can actually source data for — the data-sources section right after this one tells you where each sub-area’s data typically comes from.
Machine learning / predictive modelling
- Can a supervised model predict student academic performance from continuous-assessment scores and attendance records collected from one department?
- Which classification algorithm best predicts loan default risk using anonymised microfinance transaction data?
- Can a regression model forecast fuel-pump queue length from historical sales-volume records at one filling station chain?
- Does adding weather variables improve a model’s prediction of malaria case counts at a primary health centre, compared with case-count history alone?
- Can an ensemble model detect fraudulent mobile-money transactions more accurately than a single decision-tree model, on a labelled synthetic transaction set?
Natural language processing for Nigerian languages

- Can a sentiment-analysis model trained on Nigerian Pidgin social-media posts classify opinion polarity as accurately as one trained on standard English?
- Does a rule-based or statistical part-of-speech tagger built for Yoruba outperform an off-the-shelf English tagger applied to transliterated Yoruba text?
- Can a text-classification model route Nigerian customer-service complaints (in English and Pidgin) to the correct department more accurately than keyword matching?
- Can a small Hausa speech-to-text prototype, trained on a self-collected audio sample, transcribe short spoken commands with usable accuracy?
- Does fine-tuning an existing multilingual language model on Igbo news text improve its performance on an Igbo text-summarisation task?
Computer vision
- Can a convolutional neural network trained on a small labelled image set distinguish common cassava-leaf diseases with usable accuracy for a farmer-facing tool?
- Does a lightweight object-detection model correctly count vehicles in traffic-camera footage from one Lagos junction well enough to estimate congestion?
- Can an image-classification model distinguish counterfeit from genuine currency notes using a self-collected photo set, and how does accuracy hold up under poor lighting?
- Can a facial-recognition-adjacent attendance system correctly mark student attendance from classroom photographs without misidentifying similar-looking students?
- Does a defect-detection model trained on product images improve on manual visual inspection accuracy for one small-scale manufacturing process?
Data analytics and business intelligence
- What do two years of anonymised sales records reveal about seasonal demand patterns for one Nigerian FMCG distributor, and can a dashboard make that pattern actionable?
- Can customer-segmentation clustering on a retail loyalty-programme dataset identify a distinct high-value customer group worth a separate marketing strategy?
- Does a churn-prediction model built on telecom-adjacent usage data (call frequency, data-plan renewal history) identify at-risk subscribers earlier than a simple inactivity rule?
- What patterns does exploratory data analysis of one hospital’s anonymised outpatient records reveal about peak attendance days and staffing gaps?
Recommender systems
- Can a collaborative-filtering recommender suggest relevant courses to students based on past module choices and grades, evaluated on a held-out set of real enrolment records?
- Does a content-based recommender for an e-commerce catalogue produce more relevant product suggestions than a simple “most popular” list, measured on click-through data?
- Can a hybrid recommender improve book or article suggestions for a university library’s digital catalogue over a rules-based “same category” approach?
Time-series and forecasting
- Which forecasting model most accurately predicts short-term fuel-price movement using published historical pump-price data?
- Can a time-series model forecast weekly market prices for one staple crop (e.g. tomatoes) from historical price-bulletin data well enough to be useful to traders?
- Does an ARIMA or Prophet-based model forecast a university’s monthly electricity consumption from historical billing data more accurately than a simple moving average?
Big data and data engineering
- Can a simple ETL pipeline, built with open-source tools, clean and consolidate transaction records from three different anonymised sources into one analysis-ready dataset?
- What performance difference does a columnar storage format make when running aggregate queries on a large synthetic sales dataset, compared with a standard row-based database?
- Can a streaming data pipeline flag unusually large mobile-money transactions in near real time on a simulated transaction feed?
AI ethics, fairness and responsible use
- Does a credit-scoring model trained on historical loan data show measurably different approval rates across gender when tested on a held-out sample, and what does that imply for deployment?
- What do Nigerian bank customers report about their comfort with an AI system making loan decisions, surveyed through a structured questionnaire?
- Can a documented bias-audit checklist, applied to a student-built classification model, surface fairness issues that accuracy alone does not reveal?
- What do final year computer science students at one university report about their own use of AI tools in coursework, and does it match their department’s stated expectations?
Where Will You Get the Data for These Topics?
Every topic above lives or dies on the dataset behind it. Kaggle and the UCI Machine Learning Repository hold ready-made labelled datasets for many of the machine-learning and computer-vision topics, though few are Nigeria-specific — a project that uses one of these still needs a paragraph justifying why the dataset is a reasonable stand-in if a Nigerian-specific one is not available. For the analytics and forecasting topics, check whether a Nigerian public body such as the National Bureau of Statistics publishes the series you need, and cite only a dataset you have actually downloaded and opened, with the date you accessed it. For anything involving real transactions, patient records, or student data, the realistic path is a self-collected or anonymised institutional dataset with the relevant permission obtained first — a supervisor will ask this question at proposal stage, so answer it before you commit to a topic, not after.
Which Topics Should You Avoid?
| Topic pattern | Why it gets rejected |
|---|---|
| “Using AI to predict” some broad outcome, with no named dataset | Cannot be scoped, sourced or defended — the panel will ask “which data, which model” and there is no answer |
| A pre-trained model applied with no adaptation or evaluation | Downloading a model and running it once is not a research contribution; you need training, tuning or evaluation of your own |
| A topic requiring data you cannot legally or practically obtain (e.g. real unredacted patient records) | Ethics and data-access barriers will stall the project past your deadline |
| A generic “chatbot for” project with no dataset or evaluation plan | Heavily oversaturated and rarely includes a measurable evaluation component |
How Should You Report Results for These Topics in Chapter Four?

Whichever sub-area you pick, Chapter Four for a data science or AI project follows a different shape from a survey-based Chapter Four: instead of frequency tables and mean scores, your panel expects a stated evaluation metric appropriate to your task (accuracy, precision, recall and F1 for classification; RMSE or MAE for regression and forecasting; a confusion matrix for anything predicting categories), a comparison against at least one baseline (a simpler model, a random guess, or the previous best-known approach), and an honest account of where the model performed worst, not just its headline number. A model that scores 92 percent accuracy but fails badly on one particular subgroup or edge case is a more interesting, more defensible finding than a flat 92 percent presented with no further breakdown — “where did it fail” is a natural defence question, so plan for that answer while you are still building the model, not the night before your viva.
Frequently Asked Questions
Is Data Science and AI a separate final year project track from Computer Science in Nigerian universities?
It depends on the department — some universities run Data Science or Artificial Intelligence as a specialisation or elective track within Computer Science, while others treat it as a general project theme any computer science student can choose. Confirm with your department which topics your specific programme structure allows.
Do I need a large dataset for these projects to be accepted?
No — a small, well-documented dataset with a clearly stated limitation is more defensible than an oversized dataset you cannot explain or justify. Panels are more interested in whether you understand your data than in its size.
Can I use a publicly available international dataset instead of Nigerian data?
Yes, provided you explicitly justify the choice in your proposal and discuss in your limitations section how using non-Nigerian data affects how far your findings can generalise.
What programming tools do most Nigerian data science and AI projects use?
Python with libraries such as pandas, scikit-learn and TensorFlow or PyTorch is the most common stack, alongside Jupyter notebooks for exploratory analysis — but always confirm your department’s specific requirements or restrictions before committing.
Do I need ethics approval to use anonymised institutional data?
Most departments still require you to state how the data was anonymised and to obtain permission from whoever holds the original records, even if no individual can be identified — check this with your supervisor before you start collecting or requesting data.
Can two students in the same class both work on machine learning topics?
Yes, as long as each uses a different dataset, technique or specific research question — the sub-area is a category with dozens of distinct topics, not a single reserved topic.
Is it acceptable to fine-tune an existing pretrained model rather than building one from scratch?
Yes — this is standard practice in real data science work. What matters is that you document the fine-tuning process, your evaluation method and your results clearly, rather than presenting an unmodified pretrained model as your own contribution.
How do I decide between a machine learning topic and a data analytics topic if I am unsure which fits my skills?
A data analytics topic (exploratory analysis, dashboards, clustering) generally needs less coding depth than a full model-training machine learning topic — if you are new to Python, an analytics-first topic lets you build real skills while still meeting the project’s technical bar.
Scoping the right Data Science or AI topic is only the first decision — writing the proposal, methodology and results chapters around it is where most of the time goes. More than 9,000 students have used Tesify to write over 15,000 chapters, and every one of those chapters is 100% written by you: your dataset, your model and your analysis stay your own work. Start your project with Tesify and work through Chapters One to Five with a clear structure.
Once you have a topic, see where computer science project students in Nigeria find data for the sources not covered above, check what belongs in the limitations section of a computer science project once your model is built, read how to write up a finished software or model project as a thesis, compare system development methodologies if your project has a build component, and if you are still at proposal stage, see how to write a computer science project proposal your supervisor will approve.
