{"id":1450,"date":"2026-01-20T12:04:31","date_gmt":"2026-01-20T11:04:31","guid":{"rendered":"https:\/\/giulia-governatori.alwaysdata.net\/?p=1450"},"modified":"2026-01-20T12:11:02","modified_gmt":"2026-01-20T11:11:02","slug":"i-tried-to-use-chatgpts-architecture-to-predict-electricity-consumption-heres-what-happened","status":"publish","type":"post","link":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/i-tried-to-use-chatgpts-architecture-to-predict-electricity-consumption-heres-what-happened\/","title":{"rendered":"J'ai voulu utiliser l'architecture de ChatGPT pour pr\u00e9dire la consommation \u00e9lectrique - voici ce qui s'est pass\u00e9"},"content":{"rendered":"\n<p>Like many beginner data scientists, I was convinced that recent technologies inevitably outperformed older ones. Transformers, the neural networks powering ChatGPT that have been revolutionizing artificial intelligence since 2017, should logically crush LSTMs, an architecture invented in 1997. Twenty years apart, billions of dollars in investment, thousands of scientific publications \u2014 the match seemed decided before it started. In my case, things didn&#8217;t go as expected.<\/p>\n\n\n<style>.wp-block-kadence-advancedheading.kt-adv-heading1450_2bb75d-c7, .wp-block-kadence-advancedheading.kt-adv-heading1450_2bb75d-c7[data-kb-block=\"kb-adv-heading1450_2bb75d-c7\"]{font-style:normal;}.wp-block-kadence-advancedheading.kt-adv-heading1450_2bb75d-c7 mark.kt-highlight, .wp-block-kadence-advancedheading.kt-adv-heading1450_2bb75d-c7[data-kb-block=\"kb-adv-heading1450_2bb75d-c7\"] mark.kt-highlight{font-style:normal;color:#f76a0c;-webkit-box-decoration-break:clone;box-decoration-break:clone;padding-top:0px;padding-right:0px;padding-bottom:0px;padding-left:0px;}.wp-block-kadence-advancedheading.kt-adv-heading1450_2bb75d-c7 img.kb-inline-image, .wp-block-kadence-advancedheading.kt-adv-heading1450_2bb75d-c7[data-kb-block=\"kb-adv-heading1450_2bb75d-c7\"] img.kb-inline-image{width:150px;vertical-align:baseline;}<\/style>\n<h2 class=\"kt-adv-heading1450_2bb75d-c7 wp-block-kadence-advancedheading\" data-kb-block=\"kb-adv-heading1450_2bb75d-c7\">The challenge: anticipating consumption to optimize energy purchases<\/h2>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"alignleft size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"945\" height=\"591\" src=\"https:\/\/giulia-governatori.alwaysdata.net\/wp-content\/uploads\/1-1.webp\" alt=\"\" class=\"wp-image-1454\" style=\"width:307px;height:auto\" srcset=\"https:\/\/giulia-governatori.alwaysdata.net\/wp-content\/uploads\/1-1.webp 945w, https:\/\/giulia-governatori.alwaysdata.net\/wp-content\/uploads\/1-1-600x375.webp 600w, https:\/\/giulia-governatori.alwaysdata.net\/wp-content\/uploads\/1-1-150x94.webp 150w, https:\/\/giulia-governatori.alwaysdata.net\/wp-content\/uploads\/1-1-768x480.webp 768w, https:\/\/giulia-governatori.alwaysdata.net\/wp-content\/uploads\/1-1-18x12.webp 18w\" sizes=\"auto, (max-width: 945px) 100vw, 945px\" \/><\/figure>\n<\/div>\n\n\n<p>For my very first deep learning project, I worked on a concrete use case: predicting a household&#8217;s electricity consumption 24 hours ahead. Behind this technical problem lies a major economic challenge. Smart grid operators buy electricity daily on the European spot market (EPEX) for next-day delivery. A forecast that&#8217;s too high means waste. Too low means penalties. In my project&#8217;s fictional scenario, these errors represented 62 million euros in annual losses.<\/p>\n\n\n\n<p>I used a real dataset from UCI: four years of minute-by-minute measurements from a household in Sceaux, a Paris suburb. After hourly aggregation and feature engineering \u2014 weather variables, cyclical encoding of hours, consumption lags \u2014 I had 34,000 hours of data and 35 predictive variables.<\/p>\n\n\n<style>.wp-block-kadence-advancedheading.kt-adv-heading1450_9517d4-c3, .wp-block-kadence-advancedheading.kt-adv-heading1450_9517d4-c3[data-kb-block=\"kb-adv-heading1450_9517d4-c3\"]{font-style:normal;}.wp-block-kadence-advancedheading.kt-adv-heading1450_9517d4-c3 mark.kt-highlight, .wp-block-kadence-advancedheading.kt-adv-heading1450_9517d4-c3[data-kb-block=\"kb-adv-heading1450_9517d4-c3\"] mark.kt-highlight{font-style:normal;color:#f76a0c;-webkit-box-decoration-break:clone;box-decoration-break:clone;padding-top:0px;padding-right:0px;padding-bottom:0px;padding-left:0px;}.wp-block-kadence-advancedheading.kt-adv-heading1450_9517d4-c3 img.kb-inline-image, .wp-block-kadence-advancedheading.kt-adv-heading1450_9517d4-c3[data-kb-block=\"kb-adv-heading1450_9517d4-c3\"] img.kb-inline-image{width:150px;vertical-align:baseline;}<\/style>\n<h2 class=\"kt-adv-heading1450_9517d4-c3 wp-block-kadence-advancedheading\" data-kb-block=\"kb-adv-heading1450_9517d4-c3\">The showdown: LSTM versus Transformer<\/h2>\n\n\n\n<p>I built two models. First, the LSTM: two layers, 38,000 parameters, a classic but proven architecture. Then the Transformer: attention mechanism, positional encoding, 70,000 parameters \u2014 the heavy artillery.<\/p>\n\n\n\n<p>Raw results seemed to favor the Transformer. Its mean absolute error (MAE) reached 0.4086 kW versus 0.4145 kW for the LSTM. A 1.4% advantage. Victory? Not so fast.<\/p>\n\n\n<style>.wp-block-kadence-advancedheading.kt-adv-heading1450_68a896-f4, .wp-block-kadence-advancedheading.kt-adv-heading1450_68a896-f4[data-kb-block=\"kb-adv-heading1450_68a896-f4\"]{font-style:normal;}.wp-block-kadence-advancedheading.kt-adv-heading1450_68a896-f4 mark.kt-highlight, .wp-block-kadence-advancedheading.kt-adv-heading1450_68a896-f4[data-kb-block=\"kb-adv-heading1450_68a896-f4\"] mark.kt-highlight{font-style:normal;color:#f76a0c;-webkit-box-decoration-break:clone;box-decoration-break:clone;padding-top:0px;padding-right:0px;padding-bottom:0px;padding-left:0px;}.wp-block-kadence-advancedheading.kt-adv-heading1450_68a896-f4 img.kb-inline-image, .wp-block-kadence-advancedheading.kt-adv-heading1450_68a896-f4[data-kb-block=\"kb-adv-heading1450_68a896-f4\"] img.kb-inline-image{width:150px;vertical-align:baseline;}<\/style>\n<h2 class=\"kt-adv-heading1450_68a896-f4 wp-block-kadence-advancedheading\" data-kb-block=\"kb-adv-heading1450_68a896-f4\">The metric that changes everything: overfitting<\/h2>\n\n\n\n<p>Digging deeper into the results, I discovered a warning sign. The gap between training and test performance \u2014 what we call percentage of overfitting ((validation_loss \u2212 train_loss)\/train_loss))\u2014 reached 6.62% for the Transformer versus only 1.63% for the LSTM. In other words, the Transformer tended to memorize the training data rather than learn the true underlying patterns.<\/p>\n\n\n\n<p>In production, facing unseen data, this behavior can be catastrophic. A model that doesn&#8217;t generalize well is a dangerous model.<\/p>\n\n\n\n<p>Moreover, the LSTM trained in under 2 minutes versus over 7 for the Transformer. Simpler, faster, more robust: the choice was clear.<\/p>\n\n\n<style>.wp-block-kadence-advancedheading.kt-adv-heading1450_3ce96e-dd, .wp-block-kadence-advancedheading.kt-adv-heading1450_3ce96e-dd[data-kb-block=\"kb-adv-heading1450_3ce96e-dd\"]{font-style:normal;}.wp-block-kadence-advancedheading.kt-adv-heading1450_3ce96e-dd mark.kt-highlight, .wp-block-kadence-advancedheading.kt-adv-heading1450_3ce96e-dd[data-kb-block=\"kb-adv-heading1450_3ce96e-dd\"] mark.kt-highlight{font-style:normal;color:#f76a0c;-webkit-box-decoration-break:clone;box-decoration-break:clone;padding-top:0px;padding-right:0px;padding-bottom:0px;padding-left:0px;}.wp-block-kadence-advancedheading.kt-adv-heading1450_3ce96e-dd img.kb-inline-image, .wp-block-kadence-advancedheading.kt-adv-heading1450_3ce96e-dd[data-kb-block=\"kb-adv-heading1450_3ce96e-dd\"] img.kb-inline-image{width:150px;vertical-align:baseline;}<\/style>\n<h2 class=\"kt-adv-heading1450_3ce96e-dd wp-block-kadence-advancedheading\" data-kb-block=\"kb-adv-heading1450_3ce96e-dd\">What I take away from this<\/h2>\n\n\n\n<p>In this specific context \u2014 a single household, four years of data, relatively regular consumption patterns \u2014 the LSTM proved more suitable. The Transformer probably needs larger datasets or more complex sequences to reach its full potential. This is actually what the scientific literature suggests: attention shines on long dependencies and large data volumes, less so on short, structured time series.<\/p>\n\n\n\n<p>My final LSTM model saves approximately 28 million euros per year by reducing forecast errors by 45%. I&#8217;m proud of this result for a first deep learning project.<\/p>\n\n\n<style>.wp-block-kadence-advancedheading.kt-adv-heading1450_2da5a2-63, .wp-block-kadence-advancedheading.kt-adv-heading1450_2da5a2-63[data-kb-block=\"kb-adv-heading1450_2da5a2-63\"]{font-style:normal;}.wp-block-kadence-advancedheading.kt-adv-heading1450_2da5a2-63 mark.kt-highlight, .wp-block-kadence-advancedheading.kt-adv-heading1450_2da5a2-63[data-kb-block=\"kb-adv-heading1450_2da5a2-63\"] mark.kt-highlight{font-style:normal;color:#f76a0c;-webkit-box-decoration-break:clone;box-decoration-break:clone;padding-top:0px;padding-right:0px;padding-bottom:0px;padding-left:0px;}.wp-block-kadence-advancedheading.kt-adv-heading1450_2da5a2-63 img.kb-inline-image, .wp-block-kadence-advancedheading.kt-adv-heading1450_2da5a2-63[data-kb-block=\"kb-adv-heading1450_2da5a2-63\"] img.kb-inline-image{width:150px;vertical-align:baseline;}<\/style>\n<h2 class=\"kt-adv-heading1450_2da5a2-63 wp-block-kadence-advancedheading\" data-kb-block=\"kb-adv-heading1450_2da5a2-63\">A call to experts<\/h2>\n\n\n\n<p>That said, I&#8217;m still a beginner in this field. If you&#8217;re an experienced professional and see areas for improvement \u2014 on the Transformer architecture, hyperparameters, or training strategy \u2014 I&#8217;d love to discuss. The complete notebook and detailed methodology are available <strong><a href=\"https:\/\/giulia-governatori.alwaysdata.net\/projects\/lstm-vs-transformer-next-day-energy-forecasting-on-smart-grid\/\" data-type=\"projects\" data-id=\"1437\">here<\/a><\/strong>. Feel free to reach out: email (giuliagovernatori@hotmail.com), post Linkedin (link).<\/p>\n\n\n\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Like many beginner data scientists, I was convinced that recent technologies inevitably outperformed older ones. Transformers, the neural networks powering ChatGPT that have been revolutionizing artificial intelligence since 2017, should logically crush LSTMs, an architecture invented in 1997. Twenty years apart, billions of dollars in investment, thousands of scientific publications \u2014 the match seemed decided&#8230;<\/p>","protected":false},"author":1,"featured_media":1451,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_kad_post_transparent":"","_kad_post_title":"","_kad_post_layout":"","_kad_post_sidebar_id":"","_kad_post_content_style":"","_kad_post_vertical_padding":"","_kad_post_feature":"","_kad_post_feature_position":"","_kad_post_header":false,"_kad_post_footer":false,"_kad_post_classname":"","footnotes":""},"categories":[1],"tags":[],"class_list":["post-1450","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-non-classe"],"_links":{"self":[{"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/posts\/1450","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/comments?post=1450"}],"version-history":[{"count":3,"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/posts\/1450\/revisions"}],"predecessor-version":[{"id":1456,"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/posts\/1450\/revisions\/1456"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/media\/1451"}],"wp:attachment":[{"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/media?parent=1450"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/categories?post=1450"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/giulia-governatori.alwaysdata.net\/fr\/wp-json\/wp\/v2\/tags?post=1450"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}