Assignment

This course is assessed based on a final assignment - a computational essay.

ImportantOffice hours

Students enrolled in MZ340R18 can book a slot to discuss the assignment via cal.com/martinfleis/office-hours with Martin or via cal.com/brazdanna with Anna. The slots are available two weeks ahead and must be booked at least a day in advance.

Computational essay

A computational essay is an essay whose narrative is supported by code and its results, which are part of the essay. Think of a Jupyter Notebook with cells corresponding to text explaining the process and its results and cells with executed code doing the computation.

TipAn example of a computational essay

One nice example of a computational essay is the Age Capsule by Dani Arribas-Bel. The code in there is a bit more advanced than what you are asked to do and you will need to include some maps and possibly tables, but you get the gist.

The essay corresponds to a range of 2,500-5,000 words. That does not mean that you have to write that many words. Since you will have to produce not only text (in English or Czech) but also code and its outputs, the following requirements are specified:

  • The approximate number of words in Markdown cells is 1,500 (the bibliography, if provided, does not count towards the word count). Try to stay within 20% margin.
  • The approximate number of maps or other graphic outputs is 5 (one output may contain more than one map and will only count as one, but it must be included in the same matplotlib object).
  • The approximate number of tables is 2, where a table is considered an output of a DataFrame automatically rendered by the Notebook.

The rest of the word count is assumed to be consumed by code.

Treat this is guidelines but do not hesitate to deviate if you think it will help the narrative.

You have two options regarding topics. The first one is to work on your data on the topic of your choice, while the second is a semi-defined task if you prefer that.

Define your task

Option one is to come up with your own idea for an essay, supported by data you either already have or can gather from openly available sources.

The requirement is to cover:

  1. Initial data exploration and visualisation
  2. Exploration of a degree of randomness of data (e.g. spatial autocorrelation, point pattern analysis)
  3. At least one other technique of your choice (clustering, interpolation, regression, prediction)

Consult the second option to get a better sense of the extent.

ImportantEvery topic needs to be approved

If you decide to define your own task, the topic and the extent need to be approved by the tutor. Essays on custom topics without prior approval will not be accepted and will be marked with 0%. Reach out via email to get an approval.

The second option is more defined but still leaves some space for your creativity.

ImportantTopic for 2026/2027 to be defined

The exact topic of the pre-defined task is yet to be determined for the academic year 2026/2027. Below is the one from past years for a reference, to give you an idea.

A Barcelona case

You will take the role of a real-world data scientist tasked to explore a dataset on the city of Barcelona (Spain) and find useful insights for a variety of decision-makers. It does not matter if you have never been to Barcelona. In fact, this will help you focus on what you can learn about the city through the data, without the influence of prior knowledge. Furthermore, the assessment will not be marked based on how much you know about Barcelona but instead on how much you can show you have learned through analysing data.

Part one

In the first part, you are asked to provide an overview of the socio-economic structure of Barcelona.

Data

Head to the Open Data BCN data service of Barcelona’s City Hall and download data reflecting two aspects (two variables) of the population structure of the city at the level of census areas (Secció censal in Catalan), find relevant geometry, and link them together.

  1. Explore the spatial distribution of the data using choropleths. Comment on the details of your maps and interpret the results.
  2. Explore the degree of spatial autocorrelation. Describe the concepts behind your approach and interpret your results.

Part two

For this one, you need to pick one of the following three options.

  1. Create a classification (clustering) of Barcelona based on your socioeconomic data and interpret the results. In the process, answer the following questions:
    • What are the main types of neighbourhoods you identify?
    • Which characteristics help you delineate this typology?
    • If you had to use this classification to target areas in most need, how would you use it? Why?
    • How is the city partitioned by your data?

The other two options share the basics:

  • Download listings for Barcelona from Inside Airbnb. You have already used Airbnb data before in the course, so you can refer to the code used there. Note that you may need to use data from the archive as the recent versions do not include information about the price of each listing.
  • Barcelona is known for its issue with Airbnb density. Visualise the data appropriately and discuss why you have taken your specific approach.
  1. Asses the distribution of Airbnbs in Barcelona
    • Are the Airbnb listings distributed equally across the city? Does it depend on the type of listing or its price?
    • Can you create a regionalisation of Barcelona census areas based on the presence of Airbnbs? What does it say about the city?
  2. Asses the relationship between the socio-economic profile of Barcelona and the presence of Airbnb.
    • Use regression techniques to asses a link between the socio-economic data from part one and the variable of your choice from the Airbnb dataset. Think of a density of listings or an average price.
    • Discuss the implications of the results. What does it mean for policy?

Submission

The submission will contain an executed Jupyter Notebook. The code needs to be reproducible. That means that all the data used in the essay need to be available online (and ideally fetched directly from the notebook but a link to a download page is also fine, although data manipulation outside of the Notebook is not allowed) or shared as part of the submission. Any additional Python packages apart from those available in the provided sds environment need to be explicitly specified on top of the notebook. However, it is not expected that you will need it.

Evaluation criteria

Use of generative AI

This course recognizes that AI-assisted tools (including large language models and code assistants) are increasingly part of contemporary computational and analytical workflows. Responsible use of AI tools is permitted in this course, subject to the guidelines below.

Permitted Use

Students may use AI tools to:

  • Clarify concepts, terminology, or methodological ideas
  • Assist with debugging code or understanding error messages
  • Explore alternative approaches to a computational task

Expectations and Limitations

  • All submitted work must reflect the student’s own understanding and decision-making.
  • AI tools may not be used to generate complete solutions, analyses, or written interpretations submitted as the student’s own work.
  • Students are responsible for verifying the correctness and appropriateness of any AI-assisted code or explanations they use.
  • Work that does not demonstrate individual understanding, even if technically correct, will be marked Failed (F grade). This demonstration will be verified during an oral part of the final exam.

Disclosure Requirement

If an AI tool is used in a substantive way (for example, generating code structure, suggesting analytical steps, or drafting text), students must acknowledge its use in a short markdown cell or code comment. The disclosure should briefly state:

  • The tool used
  • The type of assistance provided

Failure to disclose substantive AI use will be treated as a violation of academic integrity.

Rationale

The purpose of this policy is to encourage transparent, ethical, and effective engagement with AI tools while maintaining the course’s emphasis on learning, mastery, and reproducibility. AI tools should support—not replace—the development of spatial data analysis skills.

This policy is adapted from the AI policy defined by prof. Sergio Rey at San Diego State University.

Acknowledgements

The assignment structure is partially derived from A Course on Geographic Data Science by Arribas-Bel (2019), licensed under CC-BY-SA 4.0.

References

Arribas-Bel, Dani. 2019. “A Course on Geographic Data Science.” The Journal of Open Source Education 2 (14). https://doi.org/10.21105/jose.00042.