Online transmission
of the 6th Congress of Polish Statistics (Warsaw July 1-2, 2026)
Day 2 (Room B, July 2)
Session 13
Nauka dla danych – nowe źródła danych i rozwój metod
Polish-language session
Session organizer: Andrzej Miszczuk
Session Chair: Andrzej Miszczuk
Session 17
Badania statystyczne – metodologia i zastosowania
Metody ilościowe w badaniach transportu, cen i zdrowia
Polish-language session
Session organizer: Krzysztof Jajuga
Session Chair: Katarzyna Kopczewska
-
Download the presentation (docx, 20 kB)
Objective
This study aims to examine how public transport accessibility affects the labour market in Poland. To this end, geographically weighted regression (GWR) will be used to examine the spatial variability of this relationship and to identify which areas of Poland could benefit most from increased accessibility. Based on the objective of the study, the following research hypotheses were formulated: 1. GWR models better describes the examined relationship than OLS models: 2. The relationship between the transport accessibility and the unemployment rate is significant and negative.
Methods
OLS and Geographically Weighted Regression (Brunsdon, Fotheringham * Charlton (1996)) were used to examine the impact of public transport availability on unemployment levels, using data from the Local Data Bank (Polish Statistics) for 2023 at the level of individual municipalities. The dependent variable in the study is the share of registered unemployed persons in the working-age population in the municipality. The main independent variable on which the analysis focuses is the developed transport accessibility index, which takes into account number of bus stops in the municipality per square kilometre, municipal expenditure on transport per capita, number of carriers registered in the voivodeship per 100,000 inhabitants, transport network density in the voivodeship, rail network density in the poviat. A key element of the analysis will be to examine the impact of this variable on the unemployment rate. A negative relationship between accessibility and the unemployment rate is expected. In addition, control variables, both demographic and economic, as well as a division into types of municipalities (urban, urban-rural, rural) were included in the model in order to avoid the error of omitting important variables. Two GWR models will be analysed: first, using an individual transport accessibility indicator for the municipality, and second, using a variable that determines the average transport accessibility of the poviat in which the municipality is located.
Results
The three tests proposed in Leung et al. (2000a) performed on GWR models showe that local models describe the relationships between variables much better than their global counterparts. Test based on Moran`s I statistic from Leung et al. (2000b) showed that spatial approach in modelling unemployment performs better than a-spatial. The spatial distribution of the regression coefficient values for the transport accessibility variable shows significant variation in both models analysed. In each model, the median coefficient is negative, which allows us to conclude that for most municipalities, improved transport accessibility is associated with a decrease in the unemployment rate.
Conclusions
Transport accessibility proved to be a significant variable explaining the level of unemployment, although its impact was not uniform across the country. In most municipalities, the regression coefficient for accessibility was negative, confirming the expected negative relationship between transport accessibility and unemployment levels. However, this variable was statistically significant only in selected regions, mainly on the peripheries of voivodeships, especially in border regions. This suggests that improving transport accessibility may be particularly effective in reducing unemployment in isolated areas most vulnerable to transport exclusion.
Keywords
transport accessibility, unemployment rate, geographically weighted regression, transport exclusion, spatial mismatch
-
Download the presentation (docx, 18 kB)
Objective
The aim of this presentation is to discuss the properties of the Young–Balk–Mehrhoff–Dikhanov (YBMD) elementary price index, which, on the one hand, is not well recognized in the literature, but on the other hand has the potential to serve as a very good alternative to the Jevons index recommended in CPI and HICP manuals. In particular, the purpose of the paper is to identify those situations in which the statistical properties of the YBMD index outperform the corresponding properties of the Jevons index.
Methods
The research methods used to achieve this objective are: (1) an analytical approach, consisting of formal mathematical proofs of selected properties of the YBMD index, and (2) an empirical study based on data collected by statistical office interviewers (data from 207 regions in Poland concerning sugar and bicycles), as well as on web-scraped data (prices of rice, yogurt, sugar, and coffee scraped from the website of a retail chain in Poland). Within the analytical approach, formulas for the bias, variance, and mean squared error (MSE) of the sample YBMD index were derived, and its properties were examined under the assumption of log-normal prices, taking into account the correlation between base-period and current-period prices. In the empirical study, in the case of data from survey regions, 207 price observations for each of the aforementioned products were used for each month covered by the study. In the case of scraped data, depending on the product, from several to several dozen observations were collected daily, depending on the number of available EAN codes for a given product on the website of the scraped retail chain. The empirical study did not attempt to generalize the findings from the samples: instead, the focus was on identifying the relationships between the values of sample elementary price indices, including the YBMD index.
Results
The results obtained in the study clearly indicate a number of advantages associated with the use of the YBMD index at the elementary level of CPI data aggregation. In particular, the main theoretical findings are as follows: (a) under log-normal prices, the sample YBMD index is an asymptotically unbiased estimator of the population YBMD index: (b) the bias and mean squared error (MSE) of the sample YBMD index depend significantly on the direction and strength of the correlation between base and current prices: (c) it is demonstrated that there exist sets of base and current prices for which the bias and MSE of the YBMD index are always smaller than the corresponding characteristics of the Jevons index. The empirical results, in turn, indicate a close approximation between the CSWD, Jevons
Conclusions
The conclusions drawn from the analytical and empirical study support recommending the YBMD index as a very good alternative to the well-established Jevons index. This recommendation is based not only on the desirable axiomatic properties of the YBMD index, but also on its statistical properties (asymptotic unbiasedness, relatively low variance, and MSE), which is a new result in the literature. A practical implication of the analysis is also that the YBMD index—as a linear approximation of the Jevons index—is potentially less sensitive to the presence of extremely low prices compared to the Jevons index (which takes the form of a product of price relatives).
Keywords
elementary indices, Jevons index, Dikhanov index, Young–Balk–Mehrhoff–Dikhanov index, Consumer Price Index (CPI)
-
Download the presentation (pdf, 3574 kB)
Objective
The COICOP classification (Classification of Individual Consumption According to Purpose) is an international classification standard published by the United Nations in 1999. It constitutes one of the key elements of modern socio-economic statistics, used to group household consumption expenditures according to their purpose. Its primary function is to ensure a consistent and comparable description of the structure of consumption across different statistical domains, in particular price statistics, national accounts, and household budget surveys.
Methods
The first version of COICOP was introduced in the early 1990s (1993) as part of the System of National Accounts. From the outset, its structure was based on a hierarchical division of expenditures into divisions, groups, and classes corresponding to different consumption categories (e.g. food, transport, health). Over time, the classification has been increasingly widely applied, becoming a foundation for the harmonisation of consumer price statistics, including CPI and HICP indices. In response to changes in consumption patterns and the development of new types of goods and services, the classification has been progressively updated. A key milestone in this evolution was the development of a new version—COICOP2018—adopted by the United Nations Statistical Commission in 2018. The updated classification better reflects contemporary consumption patterns, including the growing importance of services and the digitalisation of products. The implementation of COICOP2018 in consumer price statistics has been one of the most significant methodological undertakings in recent years. In Poland, this change has been introduced starting with data for 2026 and forms part of a broader process of modernising official statistics at the international level.
Results
The aim of this paper is to present the implementation of COICOP2018 as a systemic process encompassing methodological, organisational, and analytical dimensions. Particular emphasis is placed on the implications of this change for retail price surveys, which constitute the primary data source for CPI and HICP indices. The paper provides a detailed analysis of the challenges related to ensuring the continuity and comparability of time series. The impact of classification changes on the weighting structure and the potential implications for the level and dynamics of inflation indices are also examined.
Conclusions
The findings indicate that the implementation of COICOP2018 should be seen not only as an obligation resulting from international harmonisation, but also as an important opportunity to modernise price statistics. The change contributes to improving the representativeness of the consumer basket, enhancing cross-domain consistency, and enabling better use of new data sources. The paper falls within the thematic area of “Statistical surveys – methodology and applications”, presenting the implementation of COICOP 2018 as an example of a complex transformation process in official statistics in response to changing market conditions and increasing user demands.
Keywords
inflation, price indices, COICOP classification
-
Download the presentation (pdf, 1424 kB)
Objective
The aim of the study is to present a new method – SDEA-INI (Spatial Data Envelopment Analysis with Independent Neighbours’ Inputs) – which integrates spatial interactions while simultaneously removing the restrictive assumption of controllability of spatial inputs within the DEA framework. The proposed approach enables the assessment of regional disparities in healthcare system efficiency across Europe, accounting for both spatial dependence and the exogeneity of resources in neighbouring regions.
Methods
The study introduces a novel DEA-based solution that incorporates both spatial interactions and the exogeneity of inputs. By applying the SDEA framework, spatial autocorrelation of inputs and outputs is explicitly included in the optimisation process. Furthermore, the SDEA methodology is extended so that the new optimisation model accounts for resources located in neighbouring regions that remain beyond the control of local decision-makers. We argue that ignoring spatial interactions or the exogeneity of neighbouring inputs in productivity analysis may lead to biased results. Monte Carlo simulations clearly demonstrate that spatial extensions of DEA outperform the classical approach under spatial dependence, while SDEA-INI performs best when inputs are non-controllable. The proposed method is applied to evaluate healthcare system efficiency using Eurostat data for 232 European NUTS 2 regions for the years 2011 and 2018. Inputs include the number of doctors and hospital beds per 100,000 inhabitants, as well as GDP per capita as a proxy for financial resources. Outputs are transformed survival indicators based on mortality rates by disease groups. A spatial weights matrix based on the three nearest neighbours is incorporated into the model.
Results
The results reveal significant spatial disparities in healthcare system efficiency across Europe, with a clear pattern of an “inefficient core” (including Germany and Austria) and a more efficient periphery (Central, Eastern, and Southern Europe). The SDEA-INI method shows greater variability than standard SDEA, reducing the smoothing effect and better reflecting real-world processes. Substantial differences in regional rankings are observed when compared with DEA and SDEA. To assess result quality, sensitivity analysis (for spatial weight matrices and expenditure measures) and statistical tests (Spearman’s rank correlation, Wilcoxon test, and Dunn’s test) were conducted.
Conclusions
Accounting for spatial interactions and the non-controllability of neighbouring resources is essential for a reliable assessment of regional efficiency. The proposed SDEA-INI method contributes significantly to the development of statistical and optimisation techniques by offering a more flexible analytical tool. The findings have practical implications for healthcare policy, particularly in terms of efficient resource allocation and interregional cooperation. The proposed approach can also be applied in other areas of regional analysis.
Keywords
regional efficiency: spatial analysis: DEA: healthcare system: SDEA-INI
Session 21
Badania statystyczne – metodologia i zastosowania
Analiza i interpretacja danych w statystyce publicznej
Polish-language session
Session organizer: Krzysztof Jajuga
Session Chair: Marek Walesiak
-
Download the presentation (docx, 19 kB)
Objective
We aim to ensure the confidentiality of information collected in censuses, particularly the National Population and Housing Census (NSP). The effective application of Statistical Disclosure Control (SDC) methods is essential to achieve an optimal trade-off between minimizing the risk of unit identification and maximizing the utility of disclosed data for potential users. The purpose of the presentation is to report results of applying SDC methods and tools to data from the 2021 census.
Methods
It follows and extends the work conducted and presented in previous years. At that time, the focus was on very limited 1-km grid data about the resident population. This time, we will focus on 1-km grid population data according to the national definition, broken down by sex, economic age groups, and binary employment status: the resident population at the level of census districts, broken down by sex, economic age groups, 10-year age groups and binary employment status: and the population according to the national definition at the level of census districts, with the same breakdowns. The main SDC method used in the study was Targeted Record Swapping (TRS), recommended by Eurostat. TRS required to define e.g. a hierarchy within data, variables used to assess similarity between rows, risk threshold, variables used to risk estimation and expected swaprate. Since the TRS method does not completely eliminate the risk of indirect identification, especial probabilistic approach was also employed. It was used when several sensitive cells had minimum size: in that case the cell to be swapped was sampled with probability proportional to the share of such cell in relevant total in the population.
Results
TRS accounts for the hierarchical structure of microdata (especially their geographic dimension) and consists in identifying groups of records with the highest disclosure risk at each level of the hierarchy and swapping values for specific levels of variables within groups of similar records based on distances determined using the values of specific variables (known as similarity variables or matching keys). Values are swapped between units at a higher level. A summary of results will include an assessment of their quality, particularly from the perspective of the trade-off between data protection and utility.
Conclusions
The obtained results which will be presented, show effectiveness and practical utility of the proposed approach for efficient protection of statistical confidentiality with maximisation of utility of released data. It will be illustrated by relevant conclusions concerning the disclosure risk and information loss occuring as a result of application of statistical disclosure control methods. We will also mention potential applications of our SDC solutions to protect output data in other statistical surveys.
Keywords
statistical disclosure control, censuses, resident population, population according to the national definition, targeted record swapping
-
Download the presentation (pdf, 721 kB)
Objective
The objective of the paper is to present the role of the Statistical Metadata System (SMS) as the central platform for managing metadata in official statistics. SMS supports the design of metadata model structures, their development, collection, versioning, and dissemination. The analysis examines the impact of SMS on the quality, consistency, and stability of statistical processes, as well as its importance for building a modern metainformation infrastructure within Statistics Poland.
Methods
SMS has been collecting metadata since 2013, storing them within metadata model structures that have been continuously adapted to evolving user needs and the requirements of IT systems operating in Statistics Poland. The paper is based on an analysis of SMS functionalities and documentation describing the mechanisms that have shaped metadata management practices over more than a decade. A functional review of key system components was conducted, including metadata approval workflows, a multi-level permissions model, metadata import and export, and the integration of SMS with statistical production systems. The study also examines how metadata are used in major statistical information products such as the Domain Knowledge Bases (DBW), the Public Services Monitoring System (SMUP), the Local Data Bank (BDL), the Programme of Statistical Surveys of Official Statistics, and the Statistics Poland Information Portal. In addition, the analysis incorporates insights from long-term system operation, highlighting both strengths and limitations related to performance, usability, and technological constraints. These observations provide a basis for identifying areas requiring modernization and for understanding how SMS has informed the design of the new Metadata Subsystem.
Results
The analysis confirms that SMS effectively supports metadata management by ensuring consistency, version control, and accessibility across multiple systems in official statistics. Workflow and permission mechanisms enable quality assurance and secure handling of metadata, while integration with publication and analytical platforms stabilizes metainformation processes and reduces the risk of inconsistencies. At the same time, growing technological constraints — including insufficient automation and an outdated architecture — increasingly limit the system’s development potential, highlighting the need for modernization and redesign.
Conclusions
SMS plays a key role in the standardization and management of metadata within Statistics Poland: however, further development requires a modern architecture and full integration with other components of the metainformation system. Experience gained from long-term operation of SMS has provided the foundation for designing the new Metadata Subsystem, intended to deliver higher performance, interoperability, process automation, and a more ergonomic and coherent working environment. The findings emphasize the strategic importance of a central metadata platform for the quality and reliability of official statistics.
Keywords
metadata: metadata lifecycle: metadata standardization: SMS: system integration
-
Download the presentation (docx, 18 kB)
Objective
The aim of this work was to describe methodological changes and their impact on the measurement of labour market activity in Poland. The analysis covers discontinuities in time series arising from changes in the definitions of labour market status and the rotation scheme, the transition to computer-assisted telephone interviewing, and the update of the reference population following the 2021 National Census. Particular attention is paid to the consequences of changes introduced in 2020–2021, encompassing both the effects of the COVID-19 pandemic and the harmonisation of social surveys within t
Methods
The analyses were conducted on sub-samples of Polish LFS individual-level data for the years 1995–2024. The methods applied include descriptive statistics, linear and logistic regression, non-parametric classification models (random forest algorithm), post-stratification, and iterative proportional fitting (IPF raking). .
Results
The key factors affecting the measurement of the population and core labour supply indicators are presented. .
Conclusions
The Polish Labour Force Survey (PL-LFS), despite measurement disturbances and limited comparability of data over time, remains a key source of information on the size and structure of labour supply in Poland. It is used both for estimating the most important labour market indicators and for in-depth analyses. The identification and quantification of the effects of methodological changes are essential for the correct interpretation of labour market trends, particularly during periods of intensive change. The use of advanced weighting procedures can partially compensate for measurement disturbances.
Keywords
Polish Labour Force Survey, population, National Census, methodological changes
-
Download the presentation (pdf, 3645 kB)
Objective
The European Innovation Scoreboard is a set of indicators measuring the innovation performance of the member states of the European Union, as well as some other countries, published for over 20 years by the European Commission. Based on these indicators, a ranking of EU countries is created. In this paper, we argue that while the EIS is useful as a database for comparative analyses of selected functions of innovation systems, it is not appropriate to draw conclusions solely from a country’s position in the ranking. In particular, it should not be used as an innovation policy evaluation tool.
Methods
The data source consists of detailed rankings from the European Innovation Scoreboard published since 2014. Based on these data, as well as basic data on countries’ levels of socio-economic development, we conduct a series of analyses to determine which variables are good predictors of a country’s position in the EIS. We begin with simple correlation and rank correlation, and then move on to machine learning techniques, including decision trees. The data source consists of detailed rankings from the European Innovation Scoreboard published since 2014. Based on these data, as well as basic data on countries’ levels of socio-economic development, we conduct a series of analyses to determine which variables are good predictors of a country’s position in the EIS. We begin with simple correlation and rank correlation, and then move on to machine learning techniques, including decision trees. The data source consists of detailed rankings from the European Innovation Scoreboard published since 2014. Based on these data, as well as basic data on countries’ levels of socio-economic development, we conduct a series of analyses to determine which variables are good predictors of a country’s position in the EIS. We begin with simple correlation and rank correlation, and then move on to machine learning techniques, including decision trees.
Results
The results indicate that a country’s position in the European Innovation Scoreboard is largely explained by a limited set of variables, including GDP per capita and total R*D expenditure as a percentage of GDP (i.e., GERD). The ranking also exhibits considerable stability over time. The results indicate that a country’s position in the European Innovation Scoreboard is largely explained by a limited set of variables, including GDP per capita and total R*D expenditure as a percentage of GDP (i.e., GERD). The ranking also exhibits considerable stability over time. The results indicate that a country’s position in the European Innovation Scoreboard is largely explained by a limited set of variables, including GDP per capita and total R*D expenditure as a percentage of GDP (i.e., GERD). The
Conclusions
In public debates, one can encounter opinions that Poland’s relatively low position in the European Innovation Scoreboard indicates the ineffectiveness of innovation policy, particularly financial support for firms’ innovation activities and tax incentives. Based on the obtained results, it can be argued that such reasoning is incorrect for at least two reasons. First, it ignores the efforts by other countries, and secondly, it ignores deeper convergence process that co-determine international country rankings.
Keywords
European Innovation Scoreboard, innovation, international rankings, machine-learning
List with patronage
Honorary patronage:
Media patronage:
