A comprehensive methodology to construct standardised datasets for Science and Technology Parks
Por favor, use este identificador para citar o enlazar este ítem:
http://hdl.handle.net/10045/144922
Título: | A comprehensive methodology to construct standardised datasets for Science and Technology Parks |
---|---|
Autor/es: | Francés, Olga | Fernández Martínez, Javier | Abreu Salas, José Ignacio | Gutiérrez, Yoan | Palomar, Manuel |
Grupo/s de investigación o GITE: | Procesamiento del Lenguaje y Sistemas de Información (GPLSI) |
Centro, Departamento o Servicio: | Universidad de Alicante. Departamento de Lenguajes y Sistemas Informáticos |
Palabras clave: | Standardised dataset | Science and Technology Park (STP) | Data science | Information retrieval | Methodologies and tools |
Fecha de publicación: | 18-jun-2024 |
Editor: | Elsevier |
Cita bibliográfica: | Data & Knowledge Engineering. 2024, 153: 102338. https://doi.org/10.1016/j.datak.2024.102338 |
Resumen: | This work presents a standardised approach to create datasets for Science and Technology Parks (STPs), facilitating future analysis of STP characteristics, trends and performance. STPs are the most representative examples of innovation ecosystems. The ETL (extraction-transformation-load) structure was adapted to a global field study of STPs. A selection stage and quality check were incorporated, and the methodology was applied to Spanish STPs. This study applies diverse techniques such as expert labelling and information extraction which uses language technologies. A novel methodology for building quality and standardised STP datasets was designed and applied to a Spanish STP case study with 49 STPs. An updatable dataset and a list of the main features impacting STPs are presented. Twenty-one (n = 21) core features were refined and selected, with fifteen of them (71.4 %) being robust enough for developing further quality analysis. The methodology presented integrates different sources with heterogeneous information that is often decentralised, disaggregated and in different formats: excel files, and unstructured information in HTML or PDF format. The existence of this updatable dataset and the defined methodology will enable powerful AI tools to be applied that focus on more sophisticated analysis, such as taxonomy, monitoring, and predictive and prescriptive analytics in the innovation ecosystems field. |
Patrocinador/es: | This research is supported by the University of Alicante, the Spanish Ministry of Science and Innovation, the Generalitat Valenciana, and the European Regional Development Fund (ERDF) through the following projects: At national level, the following projects were granted: TRIVIAL(PID2021–122263OB-C22); and CORTEX(PID2021–123956OB-I00), funded by MCIN/AEI/10.13039/501100011033 and, as appropriate, by “ERDF A way of making Europe”, by the “European Union” or by the “European Union NextGenerationEU/PRTR”. At regional level, the Generalitat Valenciana (Conselleria d'Educació, Investigació, Cultura i Esport), granted project for NL4DISMIS (CIPROM/2021/21). Moreover, it was backed by the work of two COST Actions: CA19134 - “Distributed Knowledge Graphs” and CA19142 - “Leading Platform for European Citizens, Industries, Academia and Policymakers in Media Accessibility”. |
URI: | http://hdl.handle.net/10045/144922 |
ISSN: | 0169-023X (Print) | 1872-6933 (Online) |
DOI: | 10.1016/j.datak.2024.102338 |
Idioma: | eng |
Tipo: | info:eu-repo/semantics/article |
Derechos: | © 2024 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). |
Revisión científica: | si |
Versión del editor: | https://doi.org/10.1016/j.datak.2024.102338 |
Aparece en las colecciones: | INV - GPLSI - Artículos de Revistas |
Archivos en este ítem:
Archivo | Descripción | Tamaño | Formato | |
---|---|---|---|---|
![]() | 8,59 MB | Adobe PDF | Abrir Vista previa | |
Todos los documentos en RUA están protegidos por derechos de autor. Algunos derechos reservados.