What loss function to use for imbalanced classes (using PyTorch)? Announcing the arrival of Valued Associate #679: Cesar Manara Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern) 2019 Moderator Election Q&A - Questionnaire 2019 Community Moderator Election ResultsCNN - imbalanced classes, class weights vs data augmentationCensored output data, which activation function for the output layer and which loss function to use?What is the use of torch.no_grad in pytorch?Precision recall loss functionLoss function for Hierarchical Multi-label classificationHow to use Cross Entropy loss in pytorch for binary prediction?Loss function when the output is a single probabilityWhat loss function avoids overconfidence?Loss Function for Probability RegressionLoading own train data and labels in dataloader using pytorch?

How to motivate offshore teams and trust them to deliver?

What is a Meta algorithm?

List *all* the tuples!

Gastric acid as a weapon

Disable hyphenation for an entire paragraph

How to find all the available tools in macOS terminal?

Can a non-EU citizen traveling with me come with me through the EU passport line?

Is a manifold-with-boundary with given interior and non-empty boundary essentially unique?

Did Xerox really develop the first LAN?

What LEGO pieces have "real-world" functionality?

Is it ethical to give a final exam after the professor has quit before teaching the remaining chapters of the course?

How can players work together to take actions that are otherwise impossible?

Is 1 ppb equal to 1 μg/kg?

Why don't the Weasley twins use magic outside of school if the Trace can only find the location of spells cast?

What does '1 unit of lemon juice' mean in a grandma's drink recipe?

Should I call the interviewer directly, if HR aren't responding?

Do I really need recursive chmod to restrict access to a folder?

What's the difference between `auto x = vector<int>()` and `vector<int> x`?

Diagram with tikz

What is the musical term for a note that continously plays through a melody?

do i need a schengen visa for a direct flight to amsterdam?

Stars Make Stars

Antler Helmet: Can it work?

How to recreate this effect in Photoshop?



What loss function to use for imbalanced classes (using PyTorch)?



Announcing the arrival of Valued Associate #679: Cesar Manara
Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern)
2019 Moderator Election Q&A - Questionnaire
2019 Community Moderator Election ResultsCNN - imbalanced classes, class weights vs data augmentationCensored output data, which activation function for the output layer and which loss function to use?What is the use of torch.no_grad in pytorch?Precision recall loss functionLoss function for Hierarchical Multi-label classificationHow to use Cross Entropy loss in pytorch for binary prediction?Loss function when the output is a single probabilityWhat loss function avoids overconfidence?Loss Function for Probability RegressionLoading own train data and labels in dataloader using pytorch?










5












$begingroup$


I have a dataset with 3 classes with the following items:



  • Class 1: 900 elements

  • Class 2: 15000 elements

  • Class 3: 800 elements

I need to predict class 1 and class 3, which signal important deviations from the norm. Class 2 is the default “normal” case which I don’t care about.



What kind of loss function would I use here? I was thinking of using CrossEntropyLoss, but since there is a class imbalance, this would need to be weighted I suppose? How does that work in practice? Like this (using PyTorch)?



summed = 900 + 15000 + 800
weight = torch.tensor([900, 15000, 800]) / summed
crit = nn.CrossEntropyLoss(weight=weight)


Or should the weight be inverted? i.e. 1 / weight?



Is this the right approach to begin with or are there other / better methods I could use?



Thanks










share|improve this question











$endgroup$
















    5












    $begingroup$


    I have a dataset with 3 classes with the following items:



    • Class 1: 900 elements

    • Class 2: 15000 elements

    • Class 3: 800 elements

    I need to predict class 1 and class 3, which signal important deviations from the norm. Class 2 is the default “normal” case which I don’t care about.



    What kind of loss function would I use here? I was thinking of using CrossEntropyLoss, but since there is a class imbalance, this would need to be weighted I suppose? How does that work in practice? Like this (using PyTorch)?



    summed = 900 + 15000 + 800
    weight = torch.tensor([900, 15000, 800]) / summed
    crit = nn.CrossEntropyLoss(weight=weight)


    Or should the weight be inverted? i.e. 1 / weight?



    Is this the right approach to begin with or are there other / better methods I could use?



    Thanks










    share|improve this question











    $endgroup$














      5












      5








      5


      1



      $begingroup$


      I have a dataset with 3 classes with the following items:



      • Class 1: 900 elements

      • Class 2: 15000 elements

      • Class 3: 800 elements

      I need to predict class 1 and class 3, which signal important deviations from the norm. Class 2 is the default “normal” case which I don’t care about.



      What kind of loss function would I use here? I was thinking of using CrossEntropyLoss, but since there is a class imbalance, this would need to be weighted I suppose? How does that work in practice? Like this (using PyTorch)?



      summed = 900 + 15000 + 800
      weight = torch.tensor([900, 15000, 800]) / summed
      crit = nn.CrossEntropyLoss(weight=weight)


      Or should the weight be inverted? i.e. 1 / weight?



      Is this the right approach to begin with or are there other / better methods I could use?



      Thanks










      share|improve this question











      $endgroup$




      I have a dataset with 3 classes with the following items:



      • Class 1: 900 elements

      • Class 2: 15000 elements

      • Class 3: 800 elements

      I need to predict class 1 and class 3, which signal important deviations from the norm. Class 2 is the default “normal” case which I don’t care about.



      What kind of loss function would I use here? I was thinking of using CrossEntropyLoss, but since there is a class imbalance, this would need to be weighted I suppose? How does that work in practice? Like this (using PyTorch)?



      summed = 900 + 15000 + 800
      weight = torch.tensor([900, 15000, 800]) / summed
      crit = nn.CrossEntropyLoss(weight=weight)


      Or should the weight be inverted? i.e. 1 / weight?



      Is this the right approach to begin with or are there other / better methods I could use?



      Thanks







      neural-network pytorch






      share|improve this question















      share|improve this question













      share|improve this question




      share|improve this question








      edited Apr 1 at 22:37







      Muppet

















      asked Apr 1 at 19:00









      MuppetMuppet

      1485




      1485




















          1 Answer
          1






          active

          oldest

          votes


















          4












          $begingroup$


          What kind of loss function would I use here?




          Cross-entropy is the go-to loss function for classification tasks, either balanced or imbalanced. It is the first choice when no preference is built from domain knowledge yet.




          This would need to be weighted I suppose? How does that work in practice?




          Yes. Weight of class $c$ is the size of largest class divided by the size of class $c$.



          For example, If class 1 has 900, class 2 has 15000, and class 3 has 800 samples, then their weights would be 16.67, 1.0, and 18.75 respectively.



          You can also use the smallest class as nominator, which gives 0.889, 0.053, and 1.0 respectively. This is only a re-scaling, the relative weights are the same.




          Is this the right approach to begin with or are there other / better
          methods I could use?




          Yes, this is the right approach.



          EDIT:



          Thanks to @Muppet, we can also use class over-sampling, which is equivalent to using class weights. This is accomplished by WeightedRandomSampler in PyTorch, using the same aforementioned weights.






          share|improve this answer











          $endgroup$












          • $begingroup$
            I just wanted to add that using WeightedRandomSampler from PyTorch also helped, in case someone else is looking at this.
            $endgroup$
            – Muppet
            Apr 2 at 17:40











          Your Answer








          StackExchange.ready(function()
          var channelOptions =
          tags: "".split(" "),
          id: "557"
          ;
          initTagRenderer("".split(" "), "".split(" "), channelOptions);

          StackExchange.using("externalEditor", function()
          // Have to fire editor after snippets, if snippets enabled
          if (StackExchange.settings.snippets.snippetsEnabled)
          StackExchange.using("snippets", function()
          createEditor();
          );

          else
          createEditor();

          );

          function createEditor()
          StackExchange.prepareEditor(
          heartbeatType: 'answer',
          autoActivateHeartbeat: false,
          convertImagesToLinks: false,
          noModals: true,
          showLowRepImageUploadWarning: true,
          reputationToPostImages: null,
          bindNavPrevention: true,
          postfix: "",
          imageUploader:
          brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
          contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
          allowUrls: true
          ,
          onDemand: true,
          discardSelector: ".discard-answer"
          ,immediatelyShowMarkdownHelp:true
          );



          );













          draft saved

          draft discarded


















          StackExchange.ready(
          function ()
          StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fdatascience.stackexchange.com%2fquestions%2f48369%2fwhat-loss-function-to-use-for-imbalanced-classes-using-pytorch%23new-answer', 'question_page');

          );

          Post as a guest















          Required, but never shown

























          1 Answer
          1






          active

          oldest

          votes








          1 Answer
          1






          active

          oldest

          votes









          active

          oldest

          votes






          active

          oldest

          votes









          4












          $begingroup$


          What kind of loss function would I use here?




          Cross-entropy is the go-to loss function for classification tasks, either balanced or imbalanced. It is the first choice when no preference is built from domain knowledge yet.




          This would need to be weighted I suppose? How does that work in practice?




          Yes. Weight of class $c$ is the size of largest class divided by the size of class $c$.



          For example, If class 1 has 900, class 2 has 15000, and class 3 has 800 samples, then their weights would be 16.67, 1.0, and 18.75 respectively.



          You can also use the smallest class as nominator, which gives 0.889, 0.053, and 1.0 respectively. This is only a re-scaling, the relative weights are the same.




          Is this the right approach to begin with or are there other / better
          methods I could use?




          Yes, this is the right approach.



          EDIT:



          Thanks to @Muppet, we can also use class over-sampling, which is equivalent to using class weights. This is accomplished by WeightedRandomSampler in PyTorch, using the same aforementioned weights.






          share|improve this answer











          $endgroup$












          • $begingroup$
            I just wanted to add that using WeightedRandomSampler from PyTorch also helped, in case someone else is looking at this.
            $endgroup$
            – Muppet
            Apr 2 at 17:40















          4












          $begingroup$


          What kind of loss function would I use here?




          Cross-entropy is the go-to loss function for classification tasks, either balanced or imbalanced. It is the first choice when no preference is built from domain knowledge yet.




          This would need to be weighted I suppose? How does that work in practice?




          Yes. Weight of class $c$ is the size of largest class divided by the size of class $c$.



          For example, If class 1 has 900, class 2 has 15000, and class 3 has 800 samples, then their weights would be 16.67, 1.0, and 18.75 respectively.



          You can also use the smallest class as nominator, which gives 0.889, 0.053, and 1.0 respectively. This is only a re-scaling, the relative weights are the same.




          Is this the right approach to begin with or are there other / better
          methods I could use?




          Yes, this is the right approach.



          EDIT:



          Thanks to @Muppet, we can also use class over-sampling, which is equivalent to using class weights. This is accomplished by WeightedRandomSampler in PyTorch, using the same aforementioned weights.






          share|improve this answer











          $endgroup$












          • $begingroup$
            I just wanted to add that using WeightedRandomSampler from PyTorch also helped, in case someone else is looking at this.
            $endgroup$
            – Muppet
            Apr 2 at 17:40













          4












          4








          4





          $begingroup$


          What kind of loss function would I use here?




          Cross-entropy is the go-to loss function for classification tasks, either balanced or imbalanced. It is the first choice when no preference is built from domain knowledge yet.




          This would need to be weighted I suppose? How does that work in practice?




          Yes. Weight of class $c$ is the size of largest class divided by the size of class $c$.



          For example, If class 1 has 900, class 2 has 15000, and class 3 has 800 samples, then their weights would be 16.67, 1.0, and 18.75 respectively.



          You can also use the smallest class as nominator, which gives 0.889, 0.053, and 1.0 respectively. This is only a re-scaling, the relative weights are the same.




          Is this the right approach to begin with or are there other / better
          methods I could use?




          Yes, this is the right approach.



          EDIT:



          Thanks to @Muppet, we can also use class over-sampling, which is equivalent to using class weights. This is accomplished by WeightedRandomSampler in PyTorch, using the same aforementioned weights.






          share|improve this answer











          $endgroup$




          What kind of loss function would I use here?




          Cross-entropy is the go-to loss function for classification tasks, either balanced or imbalanced. It is the first choice when no preference is built from domain knowledge yet.




          This would need to be weighted I suppose? How does that work in practice?




          Yes. Weight of class $c$ is the size of largest class divided by the size of class $c$.



          For example, If class 1 has 900, class 2 has 15000, and class 3 has 800 samples, then their weights would be 16.67, 1.0, and 18.75 respectively.



          You can also use the smallest class as nominator, which gives 0.889, 0.053, and 1.0 respectively. This is only a re-scaling, the relative weights are the same.




          Is this the right approach to begin with or are there other / better
          methods I could use?




          Yes, this is the right approach.



          EDIT:



          Thanks to @Muppet, we can also use class over-sampling, which is equivalent to using class weights. This is accomplished by WeightedRandomSampler in PyTorch, using the same aforementioned weights.







          share|improve this answer














          share|improve this answer



          share|improve this answer








          edited Apr 3 at 8:18

























          answered Apr 1 at 20:29









          EsmailianEsmailian

          3,311420




          3,311420











          • $begingroup$
            I just wanted to add that using WeightedRandomSampler from PyTorch also helped, in case someone else is looking at this.
            $endgroup$
            – Muppet
            Apr 2 at 17:40
















          • $begingroup$
            I just wanted to add that using WeightedRandomSampler from PyTorch also helped, in case someone else is looking at this.
            $endgroup$
            – Muppet
            Apr 2 at 17:40















          $begingroup$
          I just wanted to add that using WeightedRandomSampler from PyTorch also helped, in case someone else is looking at this.
          $endgroup$
          – Muppet
          Apr 2 at 17:40




          $begingroup$
          I just wanted to add that using WeightedRandomSampler from PyTorch also helped, in case someone else is looking at this.
          $endgroup$
          – Muppet
          Apr 2 at 17:40

















          draft saved

          draft discarded
















































          Thanks for contributing an answer to Data Science Stack Exchange!


          • Please be sure to answer the question. Provide details and share your research!

          But avoid


          • Asking for help, clarification, or responding to other answers.

          • Making statements based on opinion; back them up with references or personal experience.

          Use MathJax to format equations. MathJax reference.


          To learn more, see our tips on writing great answers.




          draft saved


          draft discarded














          StackExchange.ready(
          function ()
          StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fdatascience.stackexchange.com%2fquestions%2f48369%2fwhat-loss-function-to-use-for-imbalanced-classes-using-pytorch%23new-answer', 'question_page');

          );

          Post as a guest















          Required, but never shown





















































          Required, but never shown














          Required, but never shown












          Required, but never shown







          Required, but never shown

































          Required, but never shown














          Required, but never shown












          Required, but never shown







          Required, but never shown







          Popular posts from this blog

          Rank groups within a grouped sequence of TRUE/FALSE and NAGrouping functions (tapply, by, aggregate) and the *apply familyCharacters counting and subletting specific patternsWhat is the purpose of setting a key in data.table?data.table vs dplyr: can one do something well the other can't or does poorly?how to make a bar plot for a list of dataframes?How to group by unique values in a list in RPandas - Alternative to rank() function that gives unique ordinal ranks for a columnRank within group in for loop in RData transformation: from dyadic to observational data in RGetting map from purrr to work with paste0

          Are all passive ability checks floors for active ability checks?Does passive perception supersede active perception?Which skills can be used passively?Active Opposition with Free-Form Professions in Fate5E Trap/Ambush/Stealth Mechanics VS Passive Perception ConfusionInteraction between perception and stealth in obscured conditionsHow does Keen Sight affect Passive Perception?Are all d20 rolls either attacks, saves or ability checks?Can players declare that they are making a specific ability check?Can I see a Hidden creature that is not obscured at all?Can a Stealth check ever be made passively?Is this alternate version of the Observant feat balanced?What is the minimum amount of skill points per HD?

          Quoting Keynes in a lectureIs differentiated instruction permitted by universities?How to make students learn prerequisitesUnsatisfactory Instructor Evaluations: balancing of expectations of engineering studentsWhat is the difference between a “statistician”, “applied statistician”, and an academic applying advanced stats within their field?Listing in reference section, but not quotingHow to efficiently use time while preparing for a class?Graduate Admissions: Teaching Emphasisstrategies for sharing teaching information with universities I don't personally have contacts withIs there an efficient way to give a large class of students feedback about their assignments?Is it unreasonable to expect students to read the lecture notes before attending the first class?