2026-02-03 03:53:09 [scrapy.utils.log] INFO: Scrapy 2.14.1 started (bot: grabbers) 2026-02-03 03:53:09 [scrapy.utils.log] INFO: Versions: {'lxml': '6.0.2', 'libxml2': '2.14.6', 'cssselect': '1.3.0', 'parsel': '1.10.0', 'w3lib': '2.3.1', 'Twisted': '25.5.0', 'Python': '3.14.2 (main, Dec 6 2025, 14:00:34) [GCC 11.4.0]', 'pyOpenSSL': '25.3.0 (OpenSSL 3.5.4 30 Sep 2025)', 'cryptography': '46.0.3', 'Platform': 'Linux-6.12.64-87.122.amzn2023.x86_64-x86_64-with-glibc2.35'} 2026-02-03 03:53:09 [scrapy.addons] INFO: Enabled addons: [] 2026-02-03 03:53:09 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/extensions/feedexport.py:436: ScrapyDeprecationWarning: The `FEED_URI` and `FEED_FORMAT` settings have been deprecated in favor of the `FEEDS` setting. Please see the `FEEDS` setting docs for more details exporter = cls(crawler) 2026-02-03 03:53:09 [scrapy.middleware] INFO: Enabled extensions: ['scrapy.extensions.corestats.CoreStats', 'scrapy.extensions.logcount.LogCount', 'scrapy.extensions.memusage.MemoryUsage', 'scrapy.extensions.logstats.LogStats', 'grabbers.extensions.custom_fields.ExportCustomFieldsExtension', 'grabbers.extensions.kol_roles.ExportKolRolesExtension', 'grabbers.extensions.exporters.DefaultScrapedItemsExporter'] 2026-02-03 03:53:09 [scrapy.crawler] INFO: Overridden settings: {'BOT_NAME': 'grabbers', 'FEED_EXPORT_FIELDS': ['External Id', 'Parent External Id', 'Name', 'Plain text: Parent Name', 'Start date', 'Start time', 'End date', 'End time', 'Location', 'Plain text: Event Link', 'Plain text: Event Link (Excel)', 'Plain text: Speaker(s)', 'Plain text: Moderator(s)', 'Plain text: Panelist(s)', 'Plain text: Chair(s)', 'Filter: Session Formats', 'Filter: Session Types', 'Filter: Working Parties'], 'FEED_FORMAT': 'xlsx', 'LOG_FILE': '/etc/scrapyd/dbs/logs/grabbers/EBMT/d5db42ba00b311f18d0a96b00f00d7fe.log', 'LOG_LEVEL': 'INFO', 'NEWSPIDER_MODULE': 'grabbers.spiders', 'SPIDER_MODULES': ['grabbers.spiders'], 'TELNETCONSOLE_ENABLED': False, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64; rv:131.0) Gecko/20100101 ' 'Firefox/131.0'} 2026-02-03 03:53:09 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-02-03 03:53:09 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-02-03 03:53:09 [scrapy.middleware] INFO: Enabled downloader middlewares: ['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware', 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware', 'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware', 'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware', 'scrapy_fake_useragent.middleware.RandomUserAgentMiddleware', 'scrapy_fake_useragent.middleware.RetryUserAgentMiddleware', 'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware', 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware', 'scrapy.downloadermiddlewares.redirect.RedirectMiddleware', 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware', 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware', 'scrapy.downloadermiddlewares.stats.DownloaderStats'] 2026-02-03 03:53:09 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:44: ScrapyDeprecationWarning: RandomUserAgentMiddleware.process_request() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_request) 2026-02-03 03:53:09 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:47: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_response() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_response) 2026-02-03 03:53:09 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:50: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_exception() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_exception) 2026-02-03 03:53:09 [scrapy.middleware] INFO: Enabled spider middlewares: ['scrapy.spidermiddlewares.start.StartSpiderMiddleware', 'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware', 'scrapy.spidermiddlewares.referer.RefererMiddleware', 'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware', 'scrapy.spidermiddlewares.depth.DepthMiddleware'] 2026-02-03 03:53:10 [scrapy.middleware] INFO: Enabled item pipelines: ['grabbers.pipelines.CleanTextPipeline', 'grabbers.pipelines.DateTimeFormatPipeline', 'grabbers.pipelines.BoolFormatPipeline', 'grabbers.pipelines.ValidateEventNamePipeline', 'grabbers.pipelines.CountryCodePipeline', 'grabbers.pipelines.HyperlinkedEventLinkPipeline', 'grabbers.pipelines.ElasticSearchPipeline'] 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CleanTextPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: DateTimeFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: BoolFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ValidateEventNamePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CountryCodePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: HyperlinkedEventLinkPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:41: ScrapyDeprecationWarning: ElasticSearchPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.open_spider) 2026-02-03 03:53:10 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ElasticSearchPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-03 03:53:10 [scrapy.core.engine] INFO: Spider opened 2026-02-03 03:53:10 [elastic_transport.transport] INFO: HEAD https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events [status:400 duration:0.118s] 2026-02-03 03:53:10 [EBMT] ERROR: ElasticSearch client initialization failed. Make sure ElasticSearch is running and connection settings are correct. Traceback (most recent call last): File "/workspace/source/app/grabbers/pipelines/store_to_elasticsearch.py", line 68, in open_spider self._prepare_es_indexes(spider) ~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^ File "/workspace/source/app/grabbers/pipelines/store_to_elasticsearch.py", line 87, in _prepare_es_indexes if self.client.index_exists(index_name=self.index_name): ~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/source/app/grabbers/elasticsearch_client.py", line 54, in index_exists return self.client.indices.exists(index=index_name) ~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^ File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/elasticsearch/_sync/client/utils.py", line 421, in wrapped return api(*args, **kwargs) File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/elasticsearch/_sync/client/indices.py", line 1592, in exists return self.perform_request( # type: ignore[return-value] ~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ "HEAD", ^^^^^^^ ...<4 lines>... path_parts=__path_parts, ^^^^^^^^^^^^^^^^^^^^^^^^ ) ^ File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/elasticsearch/_sync/client/_base.py", line 422, in perform_request return self._client.perform_request( ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^ method, ^^^^^^^ ...<5 lines>... path_parts=path_parts, ^^^^^^^^^^^^^^^^^^^^^^ ) ^ File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/elasticsearch/_sync/client/_base.py", line 271, in perform_request response = self._perform_request( method, ...<4 lines>... otel_span=otel_span, ) File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/elasticsearch/_sync/client/_base.py", line 351, in _perform_request raise HTTP_EXCEPTIONS.get(meta.status, ApiError)( message=message, meta=meta, body=resp_body ) elasticsearch.BadRequestError: BadRequestError(400, 'None') 2026-02-03 03:53:10 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min) 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG26-1 Case presentation
External ID: P89 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG26-2 Pros and Cons of evolving cellular therapies
External ID: P620 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG26-3 Challenges of Older Persons and End of Life care in cellular therapy
External ID: P621 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG26-4 Equity, Access, and Capacity - the ethical debate
External ID: P622 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG26-5 Advocacy lens on ethics of evolving therapies
External ID: P623 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG26-6 Panel Discussion - Q&A
External ID: P624 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG05 Nurses Group Workshop | Research - Please Mind the Gap: Bridging research and practice with implementation science
External ID: 9 2026-02-03 03:53:13 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: IS15 Optimizing the Transplant Experience: Patient Journey Through Conditioning and GvHD Management​ | MEDSCAPE Education Global
Supported by an Education Grant from Medac External ID: 187 2026-02-03 03:53:14 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: PSY2-3 Caregiver burden and psychosocial distress when undergoing inpatient or at-home HCT

External ID: P102 2026-02-03 03:53:14 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: PSY1-3
Preparing patients and their families for transplant: patient perception on the design and implementation of a new "preparing for transplant" seminar External ID: P98 2026-02-03 03:53:14 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: Part I: Presentation ‘What’s New in GVHD?’ External ID: P676 2026-02-03 03:53:14 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: Part II: Panel Discussion External ID: P679 2026-02-03 03:53:14 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: IS09-3 Effectively Identifying and Treating Patients with Cytomegalovirus: A Panel Discussion External ID: P719 2026-02-03 03:53:14 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG11-4 Does gene therapy change the landscape of immune deficiencies?

External ID: P47 2026-02-03 03:53:15 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: IS21-2 Patient selection:
Lessons from real-world HCT practice External ID: P612 2026-02-03 03:53:15 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: IS21-3 Optimising the patient journey:
Learnings from clinical experience in gene therapies External ID: P613 2026-02-03 03:53:15 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: LWP1-1 Report of LWP activities & introduction Jian Jian Luan award External ID: P437 2026-02-03 03:53:15 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: SS01-1 Impaired access to cellular therapy treatment and clinical trials in LMIC: impact on women

External ID: P485 2026-02-03 03:53:15 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: SS01-3 Black Women in UK STEM Higher Education: Challenges and Change

External ID: P487 2026-02-03 03:53:16 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: SS6-4 Haploidentical hematopoietic stem cell transplantation for rare diseases (Huang XiaoJun, CN)
External ID: P400 2026-02-03 03:53:16 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: SS6-5 CAR-Ts in new geographies: opportunities and barriers – The Mexican experience
External ID: P404 2026-02-03 03:53:16 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG23-3 Psychosocial - Physical - Nutrition (TBD) External ID: P665 2026-02-03 03:53:16 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: E10-3 Rare fungal infections after HCT/CT (Fusarium, Scedosporium…) External ID: P227 2026-02-03 03:53:16 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: E10-4 Mycobacterial infections after HCT/CT (TB and non-TB) External ID: P228 2026-02-03 03:53:16 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: E10-5 Rare viral infections after HCT/CT (parvovirus, HHV8, JC virus…) External ID: P229 2026-02-03 03:53:17 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: STATS2-2
Application of machine learning to case mix stratification External ID: P557 2026-02-03 03:53:17 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: IS23 Focusing GvHD care through the eyes of providers and patients: Practical insights on ECP | Delivered by Know-GvHD.com/GvHDHub.com  - Supported through an unrestricted educational grant by Therakos External ID: 211 2026-02-03 03:53:17 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: NG16-03 IMPROVING CHRONIC GVHD CARE: DEVELOPMENT OF A CHRONIC GVHD-MODULE WITHIN AN EHEALTH FACILITATED INTEGRATED CARE MODEL: SMILE External ID: P795 2026-02-03 03:53:18 [EBMT] WARNING: HTML tags found and removed from event name. Event Name: Paed1-3 T-ALL from London perspective

External ID: P497 2026-02-03 03:53:18 [scrapy.core.engine] INFO: Closing spider (finished) 2026-02-03 03:53:19 [scrapy.extensions.feedexport] INFO: Stored xlsx feed (988 items) in: s3://usummit-prod-grabbers-bucket/EBMT/EBMT-2114_2026-02-03_03-53-01.xlsx 2026-02-03 03:53:19 [scrapy.statscollectors] INFO: Dumping Scrapy stats: {'downloader/request_bytes': 97162, 'downloader/request_count': 229, 'downloader/request_method_count/GET': 229, 'downloader/response_bytes': 777298, 'downloader/response_count': 229, 'downloader/response_status_count/200': 229, 'elapsed_time_seconds': 7.656213, 'feedexport/success_count/S3FeedStorage': 1, 'finish_reason': 'finished', 'finish_time': datetime.datetime(2026, 2, 3, 3, 53, 18, 638338, tzinfo=datetime.timezone.utc), 'httpcompression/response_bytes': 3847705, 'httpcompression/response_count': 229, 'item_scraped_count': 988, 'items_per_minute': 8468.571428571428, 'log_count/INFO': 2, 'log_count/WARNING': 29, 'memusage/max': 236347392, 'memusage/startup': 236347392, 'request_depth_max': 2, 'response_received_count': 229, 'responses_per_minute': 1962.857142857143, 'scheduler/dequeued': 229, 'scheduler/dequeued/memory': 229, 'scheduler/enqueued': 229, 'scheduler/enqueued/memory': 229, 'start_time': datetime.datetime(2026, 2, 3, 3, 53, 10, 982125, tzinfo=datetime.timezone.utc)} 2026-02-03 03:53:19 [scrapy.core.engine] INFO: Spider closed (finished)