Online transmission
of the 6th Congress of Polish Statistics (Warsaw July 1-2, 2026)
Day 2 (Room D, July 2)
Session 15
Statystyka społeczna, demografia
Metodologiczne wyzwania badań demograficznych
Polish-language session
Session organizer: Elżbieta Gołata
Session Chair: Hanna Strzelecka
-
Download the presentation (docx, 20 kB)
Objective
The aim of the paper is to identify and analyse statistical disclosure control challenges for data presented as flow matrices (e.g. permanent internal migration, commuting to work, commuting to school). The main research hypothesis states that publishing such data — in particular with zero-valued flows disclosed — poses a serious risk of revealing cross-sections intended for suppression (cell suppression). The additional release of data at higher levels of territorial disaggregation may further increase this risk.
Methods
The study was conducted on real public statistics data concerning permanent internal migration, obtained from the resources of Statistics Poland (Demografia Database). The analysis covered flows at the level of communes, districts, subregions and voivodeships, allowing for a comprehensive assessment of how disclosure risk varies with the availability of different spatial aggregation levels. Methods from the field of Statistical Disclosure Control (SDC) were applied, including the analysis of sensitive cells based on the dominance rule (n, k), the analysis of zero cells as potential carriers of information, and the examination of interdependencies between cells within the matrix. A formalisation of the zero-flow problem was proposed, identifying the mechanism by which aggregation to higher spatial levels reveals additional information about flows at lower levels of data aggregation. The problem of dangerous zeros was treated separately — cells which, although formally absent from the set of sensitive cells, may in combination with an intruder`s auxiliary knowledge lead to a breach of data confidentiality at the level of individual flows. All analyses were carried out using custom programs written in the R language.
Results
The analysis conducted on data concerning inter-communal permanent migration, available in the Demografia database, demonstrated that zeros in the flow matrix are not informationally neutral — their presence unambiguously indicates the absence of specific relations between units, which in combination with additional external data may enable the indirect identification of suppressed cells. Importantly, access to complete data at a higher level of spatial aggregation (districts, voivodeships) increases the disclosure risk: the aggregated flows for larger territorial units allow for the additional reconstruction of data at the inter-communal level.
Conclusions
The results indicate the need to develop dedicated statistical disclosure control procedures for data presented in the form of flow matrices. Such procedures should account for both the informational role of zeros and the hierarchical dependencies between spatial aggregation levels. Existing confidentiality protection standards in official statistics do not sufficiently address the specific characteristics of data presented as flow matrices. The paper outlines possible directions for further research.
Keywords
statistical disclosure control, flow matrices, zero flows, primary suppression, aggregated data utilization.
-
Download the presentation (docx, 18 kB)
Objective
The LabFam Individual Biographies project aims to enhance reproducibility and comparability in social science research by developing an open, harmonized database of family and employment biographies. It focuses on reconstructing consistent fertility, partnership, and employment trajectories across five long-running panel surveys. By integrating multiple data sources and aligning life-course information across countries, the project enables robust cross-national analyses of life-course dynamics.
Methods
The study is based on the harmonization of data from five long-running panel surveys: HILDA (Australia), SOEP (Germany), SHP (Switzerland), BHPS / UKHLS (United Kingdom), and PSID (United States). For each dataset, individual biographies were reconstructed in three domains: fertility, partnership, and employment. Data from core questionnaires, calendar modules, and retrospective components were used to recover events occurring between survey waves and prior to panel entry. The results are stored as dated spells with clearly defined starts and ends, enabling analyses of durations and transitions between states. Harmonization included standardizing variable definitions, resolving inconsistencies across sources, and prioritizing information closest in time to the observed episode. In the employment domain, a common classification of labor market statuses was applied, and overlapping information was resolved using a hierarchical ordering of states. The entire process was implemented in R as open, modular code that users can reproduce, customize by country, period, and domain, and link to other infrastructures such as CPF / CNEF. Validation procedures included comparisons with indicators derived from source surveys and external demographic and labor market statistics.
Results
The project resulted in an open data infrastructure enabling the reconstruction of comparable individual histories of fertility, partnership, and employment across five countries. Internal and external validation showed a high level of consistency between LIB-derived indicators and benchmark data, including measures of childlessness, age at childbirth, age at first marriage, and labor force participation of women and men. Minor discrepancies were mainly limited to specific cohorts and datasets, confirming the overall reliability and comparability of the harmonized biographies. The highest consistency was observed for SOEP, SHP, and BHPS / UKHLS, while larger deviations appeared primarily in PSID and in the youngest or oldest cohorts.
Conclusions
LIB reduces the costs of data preparation for comparative analyses and enhances transparency in life-course research. Its open-source and modular code ensures full reproducibility, makes harmonization decisions explicit, and allows users to adapt and extend the framework to their needs. By providing harmonized, spell-based histories, the project offers a tool that integrates family and employment perspectives in cross-national analyses and supports research on the interplay between employment uncertainty, partnership formation, childbearing, and labor market withdrawal.
Keywords
data harmonization, research reproducibility
-
Download the presentation (docx, 19 kB)
Objective
The Polish Labour Force Survey (BAEL) is facing new challenges, expectations, and requirements. These stem from changing user needs and expectations, legal requirements (e.g., the new framework regulation on European social statistics), and the circumstances of fieldwork (e.g., the COVID-19 pandemic). The paper presents the changes to the way of calculation of estimation weights in the Polish LFS implemented from 2020 to the present, driven by all the factors mentioned above.
Methods
The presented actions, aimed at addressing the challenges mentioned above, involve changes in the methodology of weighting (estimation). The primary method is a calibration – calibration methods and conditions have been changed to more appropriate ones, or calibration has been introduced where it was not previously used. Regarding individual weights, the change primarily involved a shift from a simple post-stratification to a general calibration, which allowed for the inclusion of more calibration conditions to, among other things, better meet user needs. Calibration conditions have been expanded for regional estimates. Reference week conditions have been introduced, improving representativeness over time, which is particularly important in the case of unevenly distributed in time disturbances in the interviewing process (as was the case during the pandemic). In the case of household weights, a specific problem is the consequence of accepting partial unit non-responses in the survey, i.e., situations where the refusal to participate in an individual interview concerns some household members. Household weights were previously calculated as an average of household member weights and did not allow for estimates consistent with the individual estimates. The treatment of incomplete interviews in calculating household weights has been changed, and a calibration has been introduced to ensure consistency with the individual estimates.
Results
The introduction of general calibration for calculating individual weights allowed for controlled quality of estimates for a larger number of cross-sections meeting user needs, while also avoiding sample size problems in too small post-stratification strata, which negatively impacted the stability of the weights and the quality of estimates. The use of reference week based conditions ensures equal representation of all weeks of the quarter and eliminates the sensitivity of the results to variations in sample size across weeks (potential bias). Consistency was achieved between the estimates obtained using household weights and individual estimates for the basic cross-sections, including economic activity status.
Conclusions
The methodological changes allowed us to meet the requirements of the framework regulation and Eurostat`s expectations (primarily regarding weights and estimates for households), improve the quality of results, and enhance their usefulness for user needs. The change in the calculation of household weights facilitates analyses that require consideration of relationships between household members and improves the consistency of their results. The solutions for calibration with respect to reference week (time) or for dealing with incomplete household interviews may find wider application.
Keywords
Labour Force Survey, weighting, estimation, calibration
-
Download the presentation (pdf, 1689 kB)
Objective
The aim of the study is to identify the latent factors determining Healthy Life Years (HLY) in selected European Union countries and to assess the extent to which the directions of maximum variability in the set of health determinants coincide with the directions of highest predictive power for HLY. The central research question is: does supervised dimension reduction (PLS) offer an advantage over unsupervised reduction (PCR) under conditions of strong multicollinearity and a limited number of panel observations?
Methods
Two dimension reduction methods are applied: Principal Component Regression (PCR) and Partial Least Squares Regression (PLS). PCR is an unsupervised procedure in which a spectral decomposition of the correlation matrix of explanatory variables precedes OLS estimation on the selected components. PLS, in contrast to PCR, maximises the covariance between X and Y (NIPALS algorithm), which makes it a supervised method. The sample comprises 120 panel observations: 6 countries (Austria, Denmark, Germany, Italy, Poland, Spain) × 2 sexes × 10 years (2015-2024). The dependent variable (HLY) and 10 health determinants (air pollution, education, material deprivation, social protection expenditure, population density, hospital beds, physicians, alcohol consumption, smoking, obesity) are drawn from Eurostat, the World Bank and WHO. The number of components is selected via the Kaiser criterion, Horn`s parallel analysis and Leave-One-Out Cross-Validation (LOOCV) minimising RMSEP. Sample adequacy for PCA is confirmed by KMO and Bartlett tests. The models are extended with country fixed effects and a binary COVID-19 dummy. Method comparison is based on: (a) LOOCV RMSEP, (b) a paired test on squared LOO residuals (Diebold–Mariano style), (c) the correlation between PCA loadings and PLS weights, (d) bootstrap confidence intervals (B=500). All computations are performed in R using the pls, psych, factoextra and car packages.
Results
PCR attains RMSEP=2.872 with 7 components, while PLS achieves a comparable RMSEP=2.856 with only 2 components. The paired test on squared residuals reveals no significant difference in predictive accuracy (p=0.925). The low correlation between PCA loadings and PLS weights (r=0.262) confirms that the directions of maximum variance in X do not coincide with the directions of highest predictive power. In the PCR model with fixed effects, only PC5 and PC6 are significant (not PC1–PC3, which absorb 73.9% of the variance). PCR+FE: R2=0.803 with 14 parameters: PLS+FE: R2=0.762 with 9. Bootstrap (B=500) confirmed the significance of 8 out of 10 PLS coefficients.
Conclusions
The results indicate that under strong multicollinearity and a limited n / p ratio, PLS offers a clear parsimony advantage over PCR, attaining comparable predictive accuracy with 3.5 times fewer components. The pooled panel approach eliminates the estimation instability typical of per-country analyses (n / p=1:1). Distinct profiles of HLY determinants are identified: behavioural-environmental factors dominate for men, while socio-economic factors prevail for women. The contribution to the literature is a formal comparison of PCR and PLS on panel HLY data and the demonstration that standard component selection criteria (Kaiser, Horn) may be suboptimal for predictive purposes.
Keywords
Healthy Life Years (HLY): Principal Component Regression (PCR): Partial Least Squares (PLS): dimension reduction: panel data
Session 19
Zrozumieć dane: jak GUS buduje komunikację w erze informacji
Nowa strategia komunikacji GUS
Polish-language session
Session organizer: Krzysztof Jedlak, Marek Pieniążek
Session Chair: Krzysztof Jedlak, Marek Pieniążek
Panel dyskusyjny: Nikodem Chinowski, Krzysztof Jedlak, Patrycja Sikora, Sebastian Stodolak
Session 23
Statystyka gospodarcza
Wybrane problemy statystyki gospodarczej
Polish-language session
Session organizer: Eugeniusz Gatnar
Session Chair: Eugeniusz Gatnar
-
Download the presentation (pdf, 9947 kB)
Objective
The aim of the study is to assess the impact of changes in the number of foreign nationals on the economic situation of selected counties in Poland in the context of the growing and spatially differentiated scale of migration after 2019. The analysis employs the concept of “matched counties,” based on pairing units with similar characteristics in the base year, which allows for reducing the influence of initial differences and enables more reliable comparisons over time.
Methods
The study uses data from the OBM system and the Local Data Bank of Statistics Poland (BDL GUS), including information on the socio-economic situation of counties as well as data on the number and distribution of foreign nationals in Poland. The analyses are currently ongoing and are being further developed in R, using tools for data processing, statistical analysis, and matching procedures. A key methodological element is the construction of “matched counties,” i.e., pairs of territorial units formed based on the similarity of their characteristics in 2019. The matching procedure is currently being implemented and tested. Similarity is measured using the Euclidean distance (with possible alternative specifications), with the objective of minimizing the overall distance between paired counties. Two groups of units have been preliminarily identified: 10 counties with the highest and 10 with the lowest number of foreign nationals. The final construction and validation of pairs are still in progress. This approach allows for controlling initial structural differences and provides a more precise assessment of changes over time, particularly in the period 2019–2024 (and partially 2025, depending on data availability).
Results
Preliminary findings indicate potential differences in the pace of changes in the economic situation of the analyzed counties depending on the level of foreign population presence. There are signals suggesting distinct trajectories of selected indicators between counties with high and low concentrations of foreign nationals: however, these results require further verification and more in-depth analysis. The “matched counties” approach is currently being evaluated in terms of its effectiveness in reducing the impact of initial differences and better capturing potential relationships between migration and socio-economic changes.
Conclusions
The applied approach has the potential to enable a more precise analysis of the impact of changes in the number of foreign nationals on local economies: however, at the current stage, the conclusions remain preliminary. Further work will involve extending the analysis and applying more advanced econometric methods. The final results may contribute to the literature on regional economics and migration studies, as well as inform public policy at the local level. The study is exploratory in nature and provides a foundation for further research
Keywords
spatial analysis, foreign nationals, matched counties
-
Download the presentation (pdf, 1767 kB)
Objective
The main objective of the study was to assess the relationship between the functional linkage of rural and urban-rural municipalities with larger urban centers and the level of investments in the low-emission economy, co-financed by EU funds in the periods 2007–2013 and 2014–2020. The empirical analysis was conducted to verify the research hypothesis that “municipalities functionally linked to large urban centers (located within Functional Urban Areas – FUA) exhibit a higher level of investment in the development of a low-emission economy than municipalities located outside these areas.”
Methods
The study covered all rural and urban-rural municipalities in Poland, totaling 2,175 units. Data on EU co-financed projects in the field of the low-emission economy were obtained from the individual project database of the Ministry of Funds and Regional Policy. The analysis also used data from Statistics Poland (GUS), including the Local Data Bank and the “Delimitation of Rural Areas” (DOW). This typology allows municipalities to be classified according to their functional linkages with urban centers, distinguishing between agglomeration and non-agglomeration municipalities. In the first stage, the number and value of low-emission economy projects acquired by municipalities were analyzed (in absolute terms, per capita, and per km2), according to the DOW classification and voivodeships. Descriptive statistics and statistical inference methods based on significance tests were applied. In the second stage, correspondence analysis was used to graphically represent relationships between variable categories in a two-dimensional space. The input data took the form of a contingency table including two variables: (1) the type of municipality according to its affiliation with FUA and voivodeship (32 units in total), and (2) the level of investment in the low-emission economy. The quantitative variable describing the level of investment (per capita and per area unit) was categorized into four classes based on distribution quartiles.
Results
The functional linkages of the analyzed municipalities with large urban centers significantly differentiate their investment activity in the development of the low-emission economy. Correspondence analysis revealed a clear distinction between agglomeration and non-agglomeration municipalities, confirming the key role of the functional factor. Municipalities within FUAs more often implement investments at a moderate level (low or medium), whereas outside FUAs extreme patterns dominate—from no activity to high intensity. This pattern reflects center–periphery polarization mechanisms. The analysis of dispersion measures indicates that non-agglomeration municipalities are characterized by greater variability in the level of low-emission investments.
Conclusions
Development policy should be differentiated and tailored to the type and location of municipalities. It is particularly important to strengthen the institutional capacity of non-agglomeration municipalities, especially those less active in the energy transition, through financial support, advisory services, and simplified procedures. At the same time, agglomeration municipalities should be encouraged to undertake more ambitious pro-climate actions, while continued support for peripheral areas should be maintained.
Keywords
investments, energy transition, low-emission economy, Functional Urban Areas (FUA), Delimitation of Rural Areas (DOW)
-
Download the presentation (pdf, 10575 kB)
Objective
The objective of this study is to identify the strength, direction, and temporal variability of the impact of producer and importer prices on the food price index in Poland. The analysis covers selected unprocessed plant and animal products. An additional objective is to determine during which periods food price inflation was driven by changes in producer or importer prices, as well as to assess the interrelationships and feedback loops between these price categories across different time horizons.
Methods
The study employed wavelet analysis as a tool enabling the simultaneous analysis of relationships in both time and the frequency domain. This approach allows for the identification of time-varying relationships between the dynamics of producer, importer, and consumer prices, as well as for determining the direction of price impulse propagation. In particular, continuous wavelet transform and wavelet coherence analysis were used to examine the strength of interdependence and phase shifts between the analyzed time series. The empirical data include monthly price indices for Poland from Eurostat’s “Food price monitoring tool” database (https: / / doi.org / 10.2908 / PRC_FSC_IDX) and cover the period from January 2005 to December 2025, with the number of observations varying depending on the product and data availability (from n=180 to n=240). The analysis covers selected unprocessed food products, both of plant and animal origin. The study is supplemented by a Granger causality test to confirm the results of the wavelet analysis. To ensure comparability, appropriate data preprocessing procedures were applied, including the removal of seasonal components.
Results
The results indicate significant variation in the relationships between producer, importer, and consumer prices depending on the type of product. The wavelet analysis enabled the decomposition of these relationships into short-term, medium-term, and long-term components, revealing the variability of the strength of these relationships over time. In many cases, periods of dominance by producer prices are observed, particularly in long cycles, while importer prices play a significant role in shorter time horizons. The results also indicate the presence of phase shifts, which allow for the identification of delays in the transmission of price shocks between the analyzed categories. This differentiation is particularly evident between plant and animal products.
Conclusions
The use of wavelet analysis allows for an in-depth assessment of the mechanisms underlying food price formation and the identification of time-varying relationships between producer, importer, and consumer prices. The results indicate that the price transmission process is multifaceted and depends on the time horizon and the specific characteristics of the product. Eurostat’s “Food price monitoring tool” database is a valuable source of information for this type of analysis, enabling the examination of interdependencies in a dynamic context. The results obtained can be used both in macroeconomic analyses and in the formulation of agricultural and food policies.
Keywords
wavelet analysis, inflation, HICP, food prices, price transmission
-
Download the presentation (pdf, 1859 kB)
Objective
The study aims to assess to what extent the growth of technological startups operating in Warsaw is driven by their initial characteristics (initial capital endowment and its structure, number of employees, etc.) and by intra-urban location-related factors, both absolute (e.g. infrastructure accessibility) and relative (neighbourhood composition and local context). It addresses whether the impact of location is uniform, or varies depending on the initial profile of the firm.
Methods
The analysis is based on data on technological startups located in Warsaw, combining information on their initial characteristics (e.g. employee number, size, capital structure, initial financial resources and endowments) with detailed spatial indicators relating to their intra-urban location (here, both absolute and relative factors are considered, describing the neighbourhood environment, and absolute distance to the main city`s amenities). To reduce dimensionality and identify key components describing the spatial environment, Principal Component Analysis (PCA) is applied. Next, tree-based machine learning models (XGBoost) are used to capture non-linear relationships and interactions between variables. In contrast to standard econometric approaches, this framework allows for the identification of heterogeneous effects without imposing a predefined functional form. To support interpretation, Explainable Artificial Intelligence (XAI) tools are employed, including variable importance measures and partial dependence analysis. This enables the identification of both general patterns and differentiated effects across firm types, and allows to distinguish the role of absolute versus relative location factors in shaping startup growth.
Results
The results indicate that while absolute location factors, such as access to infrastructure and the main city`s amenities, are relevant for startup growth, relative factors linked to neighbourhood characteristics and local agglomeration play a more decisive role. At the same time, these effects are not uniform. A substantial degree of heterogeneity is observed: some startups benefit from being located in dense, competitive environments, while others perform better in less saturated areas. This suggests that the impact of spatial context depends strongly on initial firm characteristics and challenges approaches based on average effects.
Conclusions
The study contributes by integrating firm-level and intra-urban spatial perspectives within a single empirical framework and by explicitly addressing heterogeneity in location effects. The findings have implications for both policy and practice, suggesting that location strategies should be tailored to firm-specific profiles rather than based on general rules. From a methodological perspective, the study demonstrates the usefulness of machine learning and XAI tools in analysing complex spatial relationships and uncovering non-linear and context-dependent effects.
Keywords
technological startups, business location, agglomeration effects, machine learning, spatial analysis
List with patronage
Honorary patronage:
Media patronage:
