Machine Learning Validation Set Announcing the arrival of Valued Associate #679: Cesar Manara Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern) 2019 Moderator Election Q&A - Questionnaire 2019 Community Moderator Election ResultsHow exactly does a validation data-set work work in machine learning?Possible Reason for low Test accuracy and high AUCIntuitive interpretation of ratios between training set scores and validation set scoresReporting test result for cross-validation with Neural NetworkClustering documents - how to evaluate results?Training score at parameter tuning lower than on hold out test set (RandomForestClassifier)What to report in the build model, asses model and evaluate results steps of CRISP-DM?Machine learning - 'train_test_split' function in scikit-learn: should I repeat it several times?Why would a validation set wear out slower than a test set?Feature Scaling and normalization in cross-validation set

How to deal with a team lead who never gives me credit?

Why do people hide their license plates in the EU?

Apollo command module space walk?

How do pianists reach extremely loud dynamics?

Is it a good idea to use CNN to classify 1D signal?

Why do we bend a book to keep it straight?

Why aren't air breathing engines used as small first stages

What would be the ideal power source for a cybernetic eye?

How to remove list items depending on predecessor in python

Amount of permutations on an NxNxN Rubik's Cube

How to find all the available tools in mac terminal?

How widely used is the term Treppenwitz? Is it something that most Germans know?

Withdrew £2800, but only £2000 shows as withdrawn on online banking; what are my obligations?

A binary hook-length formula?

What does an IRS interview request entail when called in to verify expenses for a sole proprietor small business?

How does debian/ubuntu knows a package has a updated version

What is the meaning of the new sigil in Game of Thrones Season 8 intro?

Can an alien society believe that their star system is the universe?

Is it fair for a professor to grade us on the possession of past papers?

Can a USB port passively 'listen only'?

Use second argument for optional first argument if not provided in macro

Do I really need recursive chmod to restrict access to a folder?

Can I cast Passwall to drop an enemy into a 20-foot pit?

Novel: non-telepath helps overthrow rule by telepaths



Machine Learning Validation Set



Announcing the arrival of Valued Associate #679: Cesar Manara
Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern)
2019 Moderator Election Q&A - Questionnaire
2019 Community Moderator Election ResultsHow exactly does a validation data-set work work in machine learning?Possible Reason for low Test accuracy and high AUCIntuitive interpretation of ratios between training set scores and validation set scoresReporting test result for cross-validation with Neural NetworkClustering documents - how to evaluate results?Training score at parameter tuning lower than on hold out test set (RandomForestClassifier)What to report in the build model, asses model and evaluate results steps of CRISP-DM?Machine learning - 'train_test_split' function in scikit-learn: should I repeat it several times?Why would a validation set wear out slower than a test set?Feature Scaling and normalization in cross-validation set










1












$begingroup$


I have read that validation set is used for Hyper-parameter tuning and comparing models. But, what if my algorithm/model does not have any hyperparameter? Should I use validation set at all? Because comparing models can be done using Test set also.










share|improve this question









$endgroup$











  • $begingroup$
    Does your model progress in a loop? similar to neural networks? In that case you have a different model after each iteration and validation set can be used to keep the best model (at a specific iteration). Otherwise, you have only one model and validation set has no use.
    $endgroup$
    – Esmailian
    Mar 3 at 17:15











  • $begingroup$
    What do you mean to state by 'algorithm does not have any hyperparameter?'. Can you please elaborate on your problem.
    $endgroup$
    – thanatoz
    Apr 3 at 10:34















1












$begingroup$


I have read that validation set is used for Hyper-parameter tuning and comparing models. But, what if my algorithm/model does not have any hyperparameter? Should I use validation set at all? Because comparing models can be done using Test set also.










share|improve this question









$endgroup$











  • $begingroup$
    Does your model progress in a loop? similar to neural networks? In that case you have a different model after each iteration and validation set can be used to keep the best model (at a specific iteration). Otherwise, you have only one model and validation set has no use.
    $endgroup$
    – Esmailian
    Mar 3 at 17:15











  • $begingroup$
    What do you mean to state by 'algorithm does not have any hyperparameter?'. Can you please elaborate on your problem.
    $endgroup$
    – thanatoz
    Apr 3 at 10:34













1












1








1





$begingroup$


I have read that validation set is used for Hyper-parameter tuning and comparing models. But, what if my algorithm/model does not have any hyperparameter? Should I use validation set at all? Because comparing models can be done using Test set also.










share|improve this question









$endgroup$




I have read that validation set is used for Hyper-parameter tuning and comparing models. But, what if my algorithm/model does not have any hyperparameter? Should I use validation set at all? Because comparing models can be done using Test set also.







machine-learning data-science-model






share|improve this question













share|improve this question











share|improve this question




share|improve this question










asked Mar 3 at 16:35









Rishab BamraraRishab Bamrara

61




61











  • $begingroup$
    Does your model progress in a loop? similar to neural networks? In that case you have a different model after each iteration and validation set can be used to keep the best model (at a specific iteration). Otherwise, you have only one model and validation set has no use.
    $endgroup$
    – Esmailian
    Mar 3 at 17:15











  • $begingroup$
    What do you mean to state by 'algorithm does not have any hyperparameter?'. Can you please elaborate on your problem.
    $endgroup$
    – thanatoz
    Apr 3 at 10:34
















  • $begingroup$
    Does your model progress in a loop? similar to neural networks? In that case you have a different model after each iteration and validation set can be used to keep the best model (at a specific iteration). Otherwise, you have only one model and validation set has no use.
    $endgroup$
    – Esmailian
    Mar 3 at 17:15











  • $begingroup$
    What do you mean to state by 'algorithm does not have any hyperparameter?'. Can you please elaborate on your problem.
    $endgroup$
    – thanatoz
    Apr 3 at 10:34















$begingroup$
Does your model progress in a loop? similar to neural networks? In that case you have a different model after each iteration and validation set can be used to keep the best model (at a specific iteration). Otherwise, you have only one model and validation set has no use.
$endgroup$
– Esmailian
Mar 3 at 17:15





$begingroup$
Does your model progress in a loop? similar to neural networks? In that case you have a different model after each iteration and validation set can be used to keep the best model (at a specific iteration). Otherwise, you have only one model and validation set has no use.
$endgroup$
– Esmailian
Mar 3 at 17:15













$begingroup$
What do you mean to state by 'algorithm does not have any hyperparameter?'. Can you please elaborate on your problem.
$endgroup$
– thanatoz
Apr 3 at 10:34




$begingroup$
What do you mean to state by 'algorithm does not have any hyperparameter?'. Can you please elaborate on your problem.
$endgroup$
– thanatoz
Apr 3 at 10:34










2 Answers
2






active

oldest

votes


















0












$begingroup$

The validation set is there to stop you from using the test set until you are done tuning your model. When you are done tuning, you would like to have a realistic view of how the model will perform on unseen data, which is where the test set comes into play.



But tuning the model is not only hyperparameters. It involves things like feature selection, feature engineering and aslo the choice of algorithm. Even though it seems like you are already decided on a model, you should consider alternatives as it might mot be the optimal choice.






share|improve this answer









$endgroup$




















    -1












    $begingroup$

    Comparing models cannot (or should not) be done using a test set alone. You should always have a final set of data held out to estimate your generalization error. Let’s say you compare 100 different algorithms. One will eventually perform well on the test set just due to the nature of that particular data. You need the final holdout set to get a less biased estimate.



    Comparing models can be looked at the same way as tuning hyperparameters. Think of it this way, when you are tuning hyperparameters, you are comparing models. In terms of requirements comparing random forest with 200 tress vs random forest with 500 trees is no different then comparing random forest to a neural net.






    share|improve this answer









    $endgroup$













      Your Answer








      StackExchange.ready(function()
      var channelOptions =
      tags: "".split(" "),
      id: "557"
      ;
      initTagRenderer("".split(" "), "".split(" "), channelOptions);

      StackExchange.using("externalEditor", function()
      // Have to fire editor after snippets, if snippets enabled
      if (StackExchange.settings.snippets.snippetsEnabled)
      StackExchange.using("snippets", function()
      createEditor();
      );

      else
      createEditor();

      );

      function createEditor()
      StackExchange.prepareEditor(
      heartbeatType: 'answer',
      autoActivateHeartbeat: false,
      convertImagesToLinks: false,
      noModals: true,
      showLowRepImageUploadWarning: true,
      reputationToPostImages: null,
      bindNavPrevention: true,
      postfix: "",
      imageUploader:
      brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
      contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
      allowUrls: true
      ,
      onDemand: true,
      discardSelector: ".discard-answer"
      ,immediatelyShowMarkdownHelp:true
      );



      );













      draft saved

      draft discarded


















      StackExchange.ready(
      function ()
      StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fdatascience.stackexchange.com%2fquestions%2f46586%2fmachine-learning-validation-set%23new-answer', 'question_page');

      );

      Post as a guest















      Required, but never shown

























      2 Answers
      2






      active

      oldest

      votes








      2 Answers
      2






      active

      oldest

      votes









      active

      oldest

      votes






      active

      oldest

      votes









      0












      $begingroup$

      The validation set is there to stop you from using the test set until you are done tuning your model. When you are done tuning, you would like to have a realistic view of how the model will perform on unseen data, which is where the test set comes into play.



      But tuning the model is not only hyperparameters. It involves things like feature selection, feature engineering and aslo the choice of algorithm. Even though it seems like you are already decided on a model, you should consider alternatives as it might mot be the optimal choice.






      share|improve this answer









      $endgroup$

















        0












        $begingroup$

        The validation set is there to stop you from using the test set until you are done tuning your model. When you are done tuning, you would like to have a realistic view of how the model will perform on unseen data, which is where the test set comes into play.



        But tuning the model is not only hyperparameters. It involves things like feature selection, feature engineering and aslo the choice of algorithm. Even though it seems like you are already decided on a model, you should consider alternatives as it might mot be the optimal choice.






        share|improve this answer









        $endgroup$















          0












          0








          0





          $begingroup$

          The validation set is there to stop you from using the test set until you are done tuning your model. When you are done tuning, you would like to have a realistic view of how the model will perform on unseen data, which is where the test set comes into play.



          But tuning the model is not only hyperparameters. It involves things like feature selection, feature engineering and aslo the choice of algorithm. Even though it seems like you are already decided on a model, you should consider alternatives as it might mot be the optimal choice.






          share|improve this answer









          $endgroup$



          The validation set is there to stop you from using the test set until you are done tuning your model. When you are done tuning, you would like to have a realistic view of how the model will perform on unseen data, which is where the test set comes into play.



          But tuning the model is not only hyperparameters. It involves things like feature selection, feature engineering and aslo the choice of algorithm. Even though it seems like you are already decided on a model, you should consider alternatives as it might mot be the optimal choice.







          share|improve this answer












          share|improve this answer



          share|improve this answer










          answered Mar 3 at 17:24









          user10283726user10283726

          313




          313





















              -1












              $begingroup$

              Comparing models cannot (or should not) be done using a test set alone. You should always have a final set of data held out to estimate your generalization error. Let’s say you compare 100 different algorithms. One will eventually perform well on the test set just due to the nature of that particular data. You need the final holdout set to get a less biased estimate.



              Comparing models can be looked at the same way as tuning hyperparameters. Think of it this way, when you are tuning hyperparameters, you are comparing models. In terms of requirements comparing random forest with 200 tress vs random forest with 500 trees is no different then comparing random forest to a neural net.






              share|improve this answer









              $endgroup$

















                -1












                $begingroup$

                Comparing models cannot (or should not) be done using a test set alone. You should always have a final set of data held out to estimate your generalization error. Let’s say you compare 100 different algorithms. One will eventually perform well on the test set just due to the nature of that particular data. You need the final holdout set to get a less biased estimate.



                Comparing models can be looked at the same way as tuning hyperparameters. Think of it this way, when you are tuning hyperparameters, you are comparing models. In terms of requirements comparing random forest with 200 tress vs random forest with 500 trees is no different then comparing random forest to a neural net.






                share|improve this answer









                $endgroup$















                  -1












                  -1








                  -1





                  $begingroup$

                  Comparing models cannot (or should not) be done using a test set alone. You should always have a final set of data held out to estimate your generalization error. Let’s say you compare 100 different algorithms. One will eventually perform well on the test set just due to the nature of that particular data. You need the final holdout set to get a less biased estimate.



                  Comparing models can be looked at the same way as tuning hyperparameters. Think of it this way, when you are tuning hyperparameters, you are comparing models. In terms of requirements comparing random forest with 200 tress vs random forest with 500 trees is no different then comparing random forest to a neural net.






                  share|improve this answer









                  $endgroup$



                  Comparing models cannot (or should not) be done using a test set alone. You should always have a final set of data held out to estimate your generalization error. Let’s say you compare 100 different algorithms. One will eventually perform well on the test set just due to the nature of that particular data. You need the final holdout set to get a less biased estimate.



                  Comparing models can be looked at the same way as tuning hyperparameters. Think of it this way, when you are tuning hyperparameters, you are comparing models. In terms of requirements comparing random forest with 200 tress vs random forest with 500 trees is no different then comparing random forest to a neural net.







                  share|improve this answer












                  share|improve this answer



                  share|improve this answer










                  answered Mar 4 at 1:19









                  astelastel

                  1392




                  1392



























                      draft saved

                      draft discarded
















































                      Thanks for contributing an answer to Data Science Stack Exchange!


                      • Please be sure to answer the question. Provide details and share your research!

                      But avoid


                      • Asking for help, clarification, or responding to other answers.

                      • Making statements based on opinion; back them up with references or personal experience.

                      Use MathJax to format equations. MathJax reference.


                      To learn more, see our tips on writing great answers.




                      draft saved


                      draft discarded














                      StackExchange.ready(
                      function ()
                      StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fdatascience.stackexchange.com%2fquestions%2f46586%2fmachine-learning-validation-set%23new-answer', 'question_page');

                      );

                      Post as a guest















                      Required, but never shown





















































                      Required, but never shown














                      Required, but never shown












                      Required, but never shown







                      Required, but never shown

































                      Required, but never shown














                      Required, but never shown












                      Required, but never shown







                      Required, but never shown







                      Popular posts from this blog

                      Quoting Keynes in a lectureIs differentiated instruction permitted by universities?How to make students learn prerequisitesUnsatisfactory Instructor Evaluations: balancing of expectations of engineering studentsWhat is the difference between a “statistician”, “applied statistician”, and an academic applying advanced stats within their field?Listing in reference section, but not quotingHow to efficiently use time while preparing for a class?Graduate Admissions: Teaching Emphasisstrategies for sharing teaching information with universities I don't personally have contacts withIs there an efficient way to give a large class of students feedback about their assignments?Is it unreasonable to expect students to read the lecture notes before attending the first class?

                      Rank groups within a grouped sequence of TRUE/FALSE and NAGrouping functions (tapply, by, aggregate) and the *apply familyCharacters counting and subletting specific patternsWhat is the purpose of setting a key in data.table?data.table vs dplyr: can one do something well the other can't or does poorly?how to make a bar plot for a list of dataframes?How to group by unique values in a list in RPandas - Alternative to rank() function that gives unique ordinal ranks for a columnRank within group in for loop in RData transformation: from dyadic to observational data in RGetting map from purrr to work with paste0

                      Are all passive ability checks floors for active ability checks?Does passive perception supersede active perception?Which skills can be used passively?Active Opposition with Free-Form Professions in Fate5E Trap/Ambush/Stealth Mechanics VS Passive Perception ConfusionInteraction between perception and stealth in obscured conditionsHow does Keen Sight affect Passive Perception?Are all d20 rolls either attacks, saves or ability checks?Can players declare that they are making a specific ability check?Can I see a Hidden creature that is not obscured at all?Can a Stealth check ever be made passively?Is this alternate version of the Observant feat balanced?What is the minimum amount of skill points per HD?