{"id":214,"date":"2024-05-28T02:02:43","date_gmt":"2024-05-28T02:02:43","guid":{"rendered":"https:\/\/ieee-ras.conferences.computer.org\/2024\/?page_id=214"},"modified":"2024-06-10T14:09:37","modified_gmt":"2024-06-10T14:09:37","slug":"invited_talk_yogeshvarma-_abstract","status":"publish","type":"page","link":"https:\/\/ieee-ras.conferences.computer.org\/2024\/invited_talk_yogeshvarma-_abstract\/","title":{"rendered":"Invited_talk_YogeshVarma _Abstract"},"content":{"rendered":"\t\t<div data-elementor-type=\"wp-page\" data-elementor-id=\"214\" class=\"elementor elementor-214\" data-elementor-post-type=\"page\">\n\t\t\t\t<div class=\"elementor-element elementor-element-2a4eaaa e-flex e-con-boxed e-con e-parent\" data-id=\"2a4eaaa\" data-element_type=\"container\" data-e-type=\"container\">\n\t\t\t\t\t<div class=\"e-con-inner\">\n\t\t\t\t<div class=\"elementor-element elementor-element-5fea2ef elementor-widget elementor-widget-text-editor\" data-id=\"5fea2ef\" data-element_type=\"widget\" data-e-type=\"widget\" data-widget_type=\"text-editor.default\">\n\t\t\t\t<div class=\"elementor-widget-container\">\n\t\t\t\t\t\t\t\t\t<p><strong>Session \u2013 <\/strong>Hardware Fault Management<\/p>\n<p><strong>Speaker:&nbsp; <\/strong>Yogesh Varma<\/p>\n<p><strong>Title: <\/strong><u>Towards autonomous hardware fault management<\/u><\/p>\n<p>Hardware fault management of the modern data-center fleet is a complex undertaking with immense potential for improved TCO by data-driven decision-making for improving server system serviceability and availability. The fault handling consideration vary widely by server platform hardware RAS and telemetry, deployment types, and workloads. Interactable nature of hyperscale fleet error, telemetry and usage models naturally lends themselves for improved machine learning driven hardware fault management RAS actions. However, any system data-driven action can only be as good as the input data. At the Open Compute Project (OCP) Hardware Fault Management (HWFM) project Intel is partnering with key industry stakeholders to standardize a comprehensive framework for vendor agnostic hardware fault logging. This framework will enable AI assisted hardware fault analysis and autonomous RAS actions for contemporary datacenter fault management.<\/p>\n<p>This opening talk of the special session on Hardware Fault Management will set the stage by discussing the north-star for such a data-driven hardware fault management framework. It will then discuss the key initiatives and contributions. It will then naturally segway the following session to discuss details of in-band and OOB hardware fault management framework, fleet memory fault management, RAS API standard, and GPU RAS initiatives at the Open Compute Project.<\/p>\t\t\t\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t","protected":false},"excerpt":{"rendered":"<p>Session \u2013 Hardware Fault Management Speaker:&nbsp; Yogesh Varma Title: Towards autonomous hardware fault management Hardware fault management of the modern data-center fleet is a complex undertaking with immense potential for improved TCO by data-driven decision-making for improving server system serviceability and availability. The fault handling consideration vary widely by server platform hardware RAS and telemetry, [&hellip;]<\/p>\n","protected":false},"author":4,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"elementor_canvas","meta":{"footnotes":"","_members_access_role":[],"_members_access_error":""},"class_list":["post-214","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/ieee-ras.conferences.computer.org\/2024\/wp-json\/wp\/v2\/pages\/214","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ieee-ras.conferences.computer.org\/2024\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/ieee-ras.conferences.computer.org\/2024\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/ieee-ras.conferences.computer.org\/2024\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/ieee-ras.conferences.computer.org\/2024\/wp-json\/wp\/v2\/comments?post=214"}],"version-history":[{"count":0,"href":"https:\/\/ieee-ras.conferences.computer.org\/2024\/wp-json\/wp\/v2\/pages\/214\/revisions"}],"wp:attachment":[{"href":"https:\/\/ieee-ras.conferences.computer.org\/2024\/wp-json\/wp\/v2\/media?parent=214"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}