2026-02-04 12:07:37 [scrapy.utils.log] INFO: Scrapy 2.14.1 started (bot: grabbers) 2026-02-04 12:07:37 [scrapy.utils.log] INFO: Versions: {'lxml': '6.0.2', 'libxml2': '2.14.6', 'cssselect': '1.3.0', 'parsel': '1.10.0', 'w3lib': '2.3.1', 'Twisted': '25.5.0', 'Python': '3.14.2 (main, Dec 6 2025, 14:00:34) [GCC 11.4.0]', 'pyOpenSSL': '25.3.0 (OpenSSL 3.5.4 30 Sep 2025)', 'cryptography': '46.0.3', 'Platform': 'Linux-6.12.63-84.121.amzn2023.x86_64-x86_64-with-glibc2.35'} 2026-02-04 12:07:37 [scrapy.addons] INFO: Enabled addons: [] 2026-02-04 12:07:37 [scrapy.middleware] INFO: Enabled extensions: ['scrapy.extensions.corestats.CoreStats', 'scrapy.extensions.logcount.LogCount', 'scrapy.extensions.memusage.MemoryUsage', 'scrapy.extensions.logstats.LogStats', 'grabbers.extensions.custom_fields.ExportCustomFieldsExtension', 'grabbers.extensions.kol_roles.ExportKolRolesExtension'] 2026-02-04 12:07:37 [scrapy.crawler] INFO: Overridden settings: {'BOT_NAME': 'grabbers', 'FEED_EXPORT_FIELDS': ['External Id', 'Parent External Id', 'Name', 'Plain text: Parent Name', 'Start date', 'Start time', 'End date', 'End time', 'Location', 'Plain text: Event Link', 'Plain text: Event Link (Excel)', 'Plain text: Session Number', 'Plain text: Presentation Number', 'Plain text: Description', 'Filter: CME', 'Plain text: Abstract', 'Filter: Session Type', 'Filter: Track', 'Plain text: Speaker', 'Plain text: Speaker Biography'], 'LOG_FILE': '/etc/scrapyd/dbs/logs/grabbers/WCLC 2025 ' '(Events)/157a000a01c211f180e002f5c756bee6.log', 'LOG_LEVEL': 'INFO', 'NEWSPIDER_MODULE': 'grabbers.spiders', 'SPIDER_MODULES': ['grabbers.spiders'], 'TELNETCONSOLE_ENABLED': False, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64; rv:131.0) Gecko/20100101 ' 'Firefox/131.0'} 2026-02-04 12:07:37 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-02-04 12:07:37 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-02-04 12:07:37 [scrapy.middleware] INFO: Enabled downloader middlewares: ['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware', 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware', 'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware', 'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware', 'scrapy_fake_useragent.middleware.RandomUserAgentMiddleware', 'scrapy_fake_useragent.middleware.RetryUserAgentMiddleware', 'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware', 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware', 'scrapy.downloadermiddlewares.redirect.RedirectMiddleware', 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware', 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware', 'scrapy.downloadermiddlewares.stats.DownloaderStats'] 2026-02-04 12:07:37 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:44: ScrapyDeprecationWarning: RandomUserAgentMiddleware.process_request() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_request) 2026-02-04 12:07:37 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:47: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_response() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_response) 2026-02-04 12:07:37 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:50: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_exception() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_exception) 2026-02-04 12:07:37 [scrapy.middleware] INFO: Enabled spider middlewares: ['scrapy.spidermiddlewares.start.StartSpiderMiddleware', 'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware', 'scrapy.spidermiddlewares.referer.RefererMiddleware', 'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware', 'scrapy.spidermiddlewares.depth.DepthMiddleware'] 2026-02-04 12:07:38 [scrapy.middleware] INFO: Enabled item pipelines: ['grabbers.pipelines.CleanTextPipeline', 'grabbers.pipelines.DateTimeFormatPipeline', 'grabbers.pipelines.BoolFormatPipeline', 'grabbers.pipelines.ValidateEventNamePipeline', 'grabbers.pipelines.CountryCodePipeline', 'grabbers.pipelines.HyperlinkedEventLinkPipeline', 'grabbers.pipelines.ElasticSearchPipeline'] 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CleanTextPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: DateTimeFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: BoolFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ValidateEventNamePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CountryCodePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: HyperlinkedEventLinkPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:41: ScrapyDeprecationWarning: ElasticSearchPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.open_spider) 2026-02-04 12:07:38 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ElasticSearchPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-02-04 12:07:38 [scrapy.core.engine] INFO: Spider opened 2026-02-04 12:07:38 [elastic_transport.transport] INFO: HEAD https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events [status:200 duration:0.116s] 2026-02-04 12:07:38 [WCLC 2025 (Events)] INFO: Start reindexing existing events to historical index. 2026-02-04 12:07:38 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_settings [status:200 duration:0.048s] 2026-02-04 12:07:38 [elastic_transport.transport] INFO: DELETE https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events-historical?ignore_unavailable=true [status:200 duration:0.038s] 2026-02-04 12:07:39 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_clone/congress-events-historical [status:200 duration:0.244s] 2026-02-04 12:07:39 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_settings [status:200 duration:0.047s] 2026-02-04 12:07:39 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events-historical/_settings [status:200 duration:0.040s] 2026-02-04 12:07:39 [elastic_transport.transport] INFO: POST https://elasticsearch-es-http.elasticsearch.svc:9200/_reindex?wait_for_completion=true [status:200 duration:0.483s] 2026-02-04 12:07:39 [elastic_transport.transport] INFO: POST https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_delete_by_query [status:200 duration:0.006s] 2026-02-04 12:07:39 [WCLC 2025 (Events)] INFO: Events have been reindexed. 2026-02-04 12:07:39 [elastic_transport.transport] INFO: HEAD https://elasticsearch-es-http.elasticsearch.svc:9200/congress-fields [status:200 duration:0.002s] 2026-02-04 12:07:39 [WCLC 2025 (Events)] INFO: Delete existing congress fields before saving new ones. 2026-02-04 12:07:39 [elastic_transport.transport] INFO: POST https://elasticsearch-es-http.elasticsearch.svc:9200/congress-fields/_delete_by_query [status:200 duration:0.014s] 2026-02-04 12:07:39 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min) 2026-02-04 12:07:40 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Session287_congress [status:201 duration:0.057s] 2026-02-04 12:07:40 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Session287_congress [status:200 duration:0.007s] 2026-02-04 12:07:41 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Presentation3849_congress [status:201 duration:0.063s] 2026-02-04 12:07:41 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Presentation3849_congress [status:200 duration:0.007s] 2026-02-04 12:07:41 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Presentation3850_congress [status:201 duration:0.007s] 2026-02-04 12:07:41 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Presentation3851_congress [status:201 duration:0.007s] 2026-02-04 12:07:41 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Presentation3850_congress [status:200 duration:0.007s] 2026-02-04 12:07:41 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-events/_doc/Presentation3851_congress [status:200 duration:0.006s] 2026-02-04 12:07:43 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <422 https://www.abstractsonline.com/oe3/Program/21151/Search/2/Results?pagesize=3000>: HTTP status code is not handled or not allowed 2026-02-04 12:07:43 [scrapy.core.engine] INFO: Closing spider (finished) 2026-02-04 12:07:43 [WCLC 2025 (Events)] INFO: Save the actual congress fields. 2026-02-04 12:07:43 [elastic_transport.transport] INFO: PUT https://elasticsearch-es-http.elasticsearch.svc:9200/congress-fields/_doc/wclc-2025-events [status:201 duration:0.022s] 2026-02-04 12:07:43 [scrapy.statscollectors] INFO: Dumping Scrapy stats: {'downloader/request_bytes': 8028, 'downloader/request_count': 12, 'downloader/request_method_count/GET': 11, 'downloader/request_method_count/POST': 1, 'downloader/response_bytes': 32582, 'downloader/response_count': 12, 'downloader/response_status_count/200': 10, 'downloader/response_status_count/201': 1, 'downloader/response_status_count/422': 1, 'elapsed_time_seconds': 4.051871, 'finish_reason': 'finished', 'finish_time': datetime.datetime(2026, 2, 4, 12, 7, 43, 674069, tzinfo=datetime.timezone.utc), 'httperror/response_ignored_count': 1, 'httperror/response_ignored_status_count/422': 1, 'item_scraped_count': 4, 'items_per_minute': 60.0, 'log_count/INFO': 11, 'memusage/max': 216928256, 'memusage/startup': 216928256, 'request_depth_max': 4, 'response_received_count': 12, 'responses_per_minute': 180.0, 'scheduler/dequeued': 12, 'scheduler/dequeued/memory': 12, 'scheduler/enqueued': 12, 'scheduler/enqueued/memory': 12, 'start_time': datetime.datetime(2026, 2, 4, 12, 7, 39, 622198, tzinfo=datetime.timezone.utc)} 2026-02-04 12:07:43 [scrapy.core.engine] INFO: Spider closed (finished)