Skip to content

pandas 2.0 removals leave four entry points dead #1319

Description

@EnvDroneSense

Describe the bug
Four code paths call pandas APIs removed in pandas 2.0 (.append() & .iteritems()), so they raise AttributeError on any currently supported pandas. litpop.py:640 is probably the root cause of #826. None of the four is exercised by the test suite, so CI is green.

Location Removed call
climada/entity/exposures/litpop/litpop.py:640 gdf.append(...)
climada/engine/impact_data.py:422 hit_countries.append(...)
climada/engine/calibration_opt.py:406 df_result.append(df_out, input)
climada/engine/calibration_opt.py:94 years_in_common.iteritems()

To Reproduce
Steps to reproduce the behavior/error:

  1. Confirm the APIs are absent on a supported pandas.
  2. Call LitPop.from_shape_and_countries with a GeoSeries or list shape (reaches litpop.py:640), or calib_all (reaches calibration_opt.py:406).

Code example:

import pandas as pd, geopandas as gpd
print(pd.__version__)                                  # 2.2.3
print(hasattr(pd.DataFrame, "append"))                  # False
print(hasattr(gpd.GeoDataFrame, "append"))              # False
print(hasattr(pd.Series, "iteritems"))                  # False

# litpop.py:640, in the GeoSeries/list branch of `shape`:
#     gdf = gdf.append(exp.gdf.loc[exp.gdf.geometry.within(shp)])
# -> AttributeError: 'GeoDataFrame' object has no attribute 'append'

Expected behavior
These four functions run. pyproject.toml pins pandas with no upper bound, so any supported version hits this.

Climada Version: develop @ dea46fa (6.1.1-dev)

System Information (please complete the following information):

  • Operating system and version: Windows 11 Home 10.0.26200
  • Python version: 3.12.14 | packaged by conda-forge | (main, Sep 2 2026, 23:18:20) [MSC v.1944 64 bit (AMD64)]

Additional context

The fix differs per site:

  • litpop.py:640 and calibration_opt.py:406 accumulate DataFrames; collect into a list and pd.concat once. This also removes the O(n²) copy the loop currently pays, since .append reallocated the accumulated frame on every iteration.
  • impact_data.py:422 appends a dict, not a frame; collect the rows into a list and build one pd.DataFrame(records, columns=[...]) at the end.
  • calibration_opt.py:94 needs no restructuring: Series.iteritems() -> Series.items().

calibration_opt.py:406 also passes the builtin input as the second positional argument (ignore_index). Since bool(input) is True, it happened to mean ignore_index=True, so this was a latent bug rather than a visible one.

Two adjacent observations in the litpop.py loop, neither part of the fix above:

  • exp.gdf.geometry.within(shp) scans every exposure point per shape with no spatial index, so the loop is O(n_shapes x n_points). The Shape branch at :626 already uses geopandas.sjoin for the same job.
  • A point inside two overlapping shapes is matched twice and appears twice in the result, so its value is double-counted in the exposure total. Reproduced on a 5-point toy case with two overlapping squares: 6 rows summing to 180.0, where the unique total is 150.0. geopandas.sjoin(..., predicate="within", how="inner") returns the identical row multiset, so switching would not change this. Whether it should be deduplicated is a question for the maintainers, since it would alter exposure totals for overlapping input shapes.

The test-coverage gap is arguably the more significant finding: four user-facing entry points are dead and nothing detected it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions