Multiple filtering pandas columns based on values in another columnCreating new columns by iterating over rows in pandas dataframePandas - Get feature values which appear in two distinct dataframesPandas Query Optimization On Multiple Columnshow many rows have values from the same columns pandasExport pandas to dictionary by combining multiple row valuesCombine Pandas DataFrames with year columnsSpearmanr on two pandas dataframesShould I use pandas get_dummies and create additional columns or use my own encoding code that keeps 1 column?Merging common Columns values in two DataFrame PandasAggregate values of same name pandas dataframe columns to single column

In a multiple cat home, how many litter boxes should you have?

What is going on with gets(stdin) on the site coderbyte?

US tourist/student visa

Did the UK lift the requirement for registering SIM cards?

I found an audio circuit and I built it just fine, but I find it a bit too quiet. How do I amplify the output so that it is a bit louder?

What to do when eye contact makes your coworker uncomfortable?

How to explain what's wrong with this application of the chain rule?

Why does Carol not get rid of the Kree symbol on her suit when she changes its colours?

What does "Scientists rise up against statistical significance" mean? (Comment in Nature)

Merge org tables

Why the "ls" command is showing the permissions of files in a FAT32 partition?

Is there a RAID 0 Equivalent for RAM?

Make a Bowl of Alphabet Soup

The IT department bottlenecks progress, how should I handle this?

Biological Blimps: Propulsion

"It doesn't matter" or "it won't matter"?

How much theory knowledge is actually used while playing?

The Digit Triangles

What is Cash Advance APR?

Does an advisor owe his/her student anything? Will an advisor keep a PhD student only out of pity?

Does grappling negate Mirror Image?

Are Captain Marvel's powers affected by Thanos breaking the Tesseract and claiming the stone?

How to draw a matrix with arrows in limited space

Microchip documentation does not label CAN buss pins on micro controller pinout diagram



Multiple filtering pandas columns based on values in another column


Creating new columns by iterating over rows in pandas dataframePandas - Get feature values which appear in two distinct dataframesPandas Query Optimization On Multiple Columnshow many rows have values from the same columns pandasExport pandas to dictionary by combining multiple row valuesCombine Pandas DataFrames with year columnsSpearmanr on two pandas dataframesShould I use pandas get_dummies and create additional columns or use my own encoding code that keeps 1 column?Merging common Columns values in two DataFrame PandasAggregate values of same name pandas dataframe columns to single column













0












$begingroup$


I have a pandas dataframe df1:



df1



Now, I want to filter the rows in df1 based on unique combinations of (Campaign, Merchant) from another dataframe, df2, which look like this:



enter image description here



What I tried is using .isin, with a code similar to the one below:



df1.loc[df1['Campaign'].isin(df2['Campaign']) &
df1['Merchant'].isin(df2['Merchant'])]


The problem here is that the conditions are independent eg : I want to check if (A,1) from df2 is in df1, but with the above condition, since I am checking all the list, not row by row, it would return all rows in df1 where Campaign column is A OR Merchant column is 1.



Do you have any suggestion for this multiple pandas filtering?










share|improve this question











$endgroup$
















    0












    $begingroup$


    I have a pandas dataframe df1:



    df1



    Now, I want to filter the rows in df1 based on unique combinations of (Campaign, Merchant) from another dataframe, df2, which look like this:



    enter image description here



    What I tried is using .isin, with a code similar to the one below:



    df1.loc[df1['Campaign'].isin(df2['Campaign']) &
    df1['Merchant'].isin(df2['Merchant'])]


    The problem here is that the conditions are independent eg : I want to check if (A,1) from df2 is in df1, but with the above condition, since I am checking all the list, not row by row, it would return all rows in df1 where Campaign column is A OR Merchant column is 1.



    Do you have any suggestion for this multiple pandas filtering?










    share|improve this question











    $endgroup$














      0












      0








      0





      $begingroup$


      I have a pandas dataframe df1:



      df1



      Now, I want to filter the rows in df1 based on unique combinations of (Campaign, Merchant) from another dataframe, df2, which look like this:



      enter image description here



      What I tried is using .isin, with a code similar to the one below:



      df1.loc[df1['Campaign'].isin(df2['Campaign']) &
      df1['Merchant'].isin(df2['Merchant'])]


      The problem here is that the conditions are independent eg : I want to check if (A,1) from df2 is in df1, but with the above condition, since I am checking all the list, not row by row, it would return all rows in df1 where Campaign column is A OR Merchant column is 1.



      Do you have any suggestion for this multiple pandas filtering?










      share|improve this question











      $endgroup$




      I have a pandas dataframe df1:



      df1



      Now, I want to filter the rows in df1 based on unique combinations of (Campaign, Merchant) from another dataframe, df2, which look like this:



      enter image description here



      What I tried is using .isin, with a code similar to the one below:



      df1.loc[df1['Campaign'].isin(df2['Campaign']) &
      df1['Merchant'].isin(df2['Merchant'])]


      The problem here is that the conditions are independent eg : I want to check if (A,1) from df2 is in df1, but with the above condition, since I am checking all the list, not row by row, it would return all rows in df1 where Campaign column is A OR Merchant column is 1.



      Do you have any suggestion for this multiple pandas filtering?







      python pandas






      share|improve this question















      share|improve this question













      share|improve this question




      share|improve this question








      edited 2 days ago









      tuomastik

      753418




      753418










      asked Mar 18 at 21:25









      Remus RaphaelRemus Raphael

      112




      112




















          1 Answer
          1






          active

          oldest

          votes


















          0












          $begingroup$

          import pandas as pd

          df1 = pd.DataFrame("Random numbers 1": pd.np.random.randn(6),
          "Campaign": ["A"] * 5 + ["B"],
          "Merchant": [1, 1, 1, 2, 3, 1])

          df2 = pd.DataFrame("Random numbers 2": pd.np.random.randn(6),
          "Campaign": ["A"] * 2 + ["B"] * 2 + ["C"] * 2,
          "Merchant": [1, 2, 1, 2, 1, 2])

          columns_consider = ["Campaign", "Merchant"]
          combined = pd.concat((df1[columns_consider].drop_duplicates(),
          df2[columns_consider].drop_duplicates()), ignore_index=True)

          identical = combined[combined.duplicated()]

          print(identical)


          Output:



           Campaign Merchant
          4 A 1
          5 A 2
          6 B 1





          share|improve this answer









          $endgroup$












            Your Answer





            StackExchange.ifUsing("editor", function ()
            return StackExchange.using("mathjaxEditing", function ()
            StackExchange.MarkdownEditor.creationCallbacks.add(function (editor, postfix)
            StackExchange.mathjaxEditing.prepareWmdForMathJax(editor, postfix, [["$", "$"], ["\\(","\\)"]]);
            );
            );
            , "mathjax-editing");

            StackExchange.ready(function()
            var channelOptions =
            tags: "".split(" "),
            id: "557"
            ;
            initTagRenderer("".split(" "), "".split(" "), channelOptions);

            StackExchange.using("externalEditor", function()
            // Have to fire editor after snippets, if snippets enabled
            if (StackExchange.settings.snippets.snippetsEnabled)
            StackExchange.using("snippets", function()
            createEditor();
            );

            else
            createEditor();

            );

            function createEditor()
            StackExchange.prepareEditor(
            heartbeatType: 'answer',
            autoActivateHeartbeat: false,
            convertImagesToLinks: false,
            noModals: true,
            showLowRepImageUploadWarning: true,
            reputationToPostImages: null,
            bindNavPrevention: true,
            postfix: "",
            imageUploader:
            brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
            contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
            allowUrls: true
            ,
            onDemand: true,
            discardSelector: ".discard-answer"
            ,immediatelyShowMarkdownHelp:true
            );



            );













            draft saved

            draft discarded


















            StackExchange.ready(
            function ()
            StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fdatascience.stackexchange.com%2fquestions%2f47562%2fmultiple-filtering-pandas-columns-based-on-values-in-another-column%23new-answer', 'question_page');

            );

            Post as a guest















            Required, but never shown

























            1 Answer
            1






            active

            oldest

            votes








            1 Answer
            1






            active

            oldest

            votes









            active

            oldest

            votes






            active

            oldest

            votes









            0












            $begingroup$

            import pandas as pd

            df1 = pd.DataFrame("Random numbers 1": pd.np.random.randn(6),
            "Campaign": ["A"] * 5 + ["B"],
            "Merchant": [1, 1, 1, 2, 3, 1])

            df2 = pd.DataFrame("Random numbers 2": pd.np.random.randn(6),
            "Campaign": ["A"] * 2 + ["B"] * 2 + ["C"] * 2,
            "Merchant": [1, 2, 1, 2, 1, 2])

            columns_consider = ["Campaign", "Merchant"]
            combined = pd.concat((df1[columns_consider].drop_duplicates(),
            df2[columns_consider].drop_duplicates()), ignore_index=True)

            identical = combined[combined.duplicated()]

            print(identical)


            Output:



             Campaign Merchant
            4 A 1
            5 A 2
            6 B 1





            share|improve this answer









            $endgroup$

















              0












              $begingroup$

              import pandas as pd

              df1 = pd.DataFrame("Random numbers 1": pd.np.random.randn(6),
              "Campaign": ["A"] * 5 + ["B"],
              "Merchant": [1, 1, 1, 2, 3, 1])

              df2 = pd.DataFrame("Random numbers 2": pd.np.random.randn(6),
              "Campaign": ["A"] * 2 + ["B"] * 2 + ["C"] * 2,
              "Merchant": [1, 2, 1, 2, 1, 2])

              columns_consider = ["Campaign", "Merchant"]
              combined = pd.concat((df1[columns_consider].drop_duplicates(),
              df2[columns_consider].drop_duplicates()), ignore_index=True)

              identical = combined[combined.duplicated()]

              print(identical)


              Output:



               Campaign Merchant
              4 A 1
              5 A 2
              6 B 1





              share|improve this answer









              $endgroup$















                0












                0








                0





                $begingroup$

                import pandas as pd

                df1 = pd.DataFrame("Random numbers 1": pd.np.random.randn(6),
                "Campaign": ["A"] * 5 + ["B"],
                "Merchant": [1, 1, 1, 2, 3, 1])

                df2 = pd.DataFrame("Random numbers 2": pd.np.random.randn(6),
                "Campaign": ["A"] * 2 + ["B"] * 2 + ["C"] * 2,
                "Merchant": [1, 2, 1, 2, 1, 2])

                columns_consider = ["Campaign", "Merchant"]
                combined = pd.concat((df1[columns_consider].drop_duplicates(),
                df2[columns_consider].drop_duplicates()), ignore_index=True)

                identical = combined[combined.duplicated()]

                print(identical)


                Output:



                 Campaign Merchant
                4 A 1
                5 A 2
                6 B 1





                share|improve this answer









                $endgroup$



                import pandas as pd

                df1 = pd.DataFrame("Random numbers 1": pd.np.random.randn(6),
                "Campaign": ["A"] * 5 + ["B"],
                "Merchant": [1, 1, 1, 2, 3, 1])

                df2 = pd.DataFrame("Random numbers 2": pd.np.random.randn(6),
                "Campaign": ["A"] * 2 + ["B"] * 2 + ["C"] * 2,
                "Merchant": [1, 2, 1, 2, 1, 2])

                columns_consider = ["Campaign", "Merchant"]
                combined = pd.concat((df1[columns_consider].drop_duplicates(),
                df2[columns_consider].drop_duplicates()), ignore_index=True)

                identical = combined[combined.duplicated()]

                print(identical)


                Output:



                 Campaign Merchant
                4 A 1
                5 A 2
                6 B 1






                share|improve this answer












                share|improve this answer



                share|improve this answer










                answered 2 days ago









                tuomastiktuomastik

                753418




                753418



























                    draft saved

                    draft discarded
















































                    Thanks for contributing an answer to Data Science Stack Exchange!


                    • Please be sure to answer the question. Provide details and share your research!

                    But avoid


                    • Asking for help, clarification, or responding to other answers.

                    • Making statements based on opinion; back them up with references or personal experience.

                    Use MathJax to format equations. MathJax reference.


                    To learn more, see our tips on writing great answers.




                    draft saved


                    draft discarded














                    StackExchange.ready(
                    function ()
                    StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fdatascience.stackexchange.com%2fquestions%2f47562%2fmultiple-filtering-pandas-columns-based-on-values-in-another-column%23new-answer', 'question_page');

                    );

                    Post as a guest















                    Required, but never shown





















































                    Required, but never shown














                    Required, but never shown












                    Required, but never shown







                    Required, but never shown

































                    Required, but never shown














                    Required, but never shown












                    Required, but never shown







                    Required, but never shown







                    Popular posts from this blog

                    Quoting Keynes in a lectureIs differentiated instruction permitted by universities?How to make students learn prerequisitesUnsatisfactory Instructor Evaluations: balancing of expectations of engineering studentsWhat is the difference between a “statistician”, “applied statistician”, and an academic applying advanced stats within their field?Listing in reference section, but not quotingHow to efficiently use time while preparing for a class?Graduate Admissions: Teaching Emphasisstrategies for sharing teaching information with universities I don't personally have contacts withIs there an efficient way to give a large class of students feedback about their assignments?Is it unreasonable to expect students to read the lecture notes before attending the first class?

                    Rank groups within a grouped sequence of TRUE/FALSE and NAGrouping functions (tapply, by, aggregate) and the *apply familyCharacters counting and subletting specific patternsWhat is the purpose of setting a key in data.table?data.table vs dplyr: can one do something well the other can't or does poorly?how to make a bar plot for a list of dataframes?How to group by unique values in a list in RPandas - Alternative to rank() function that gives unique ordinal ranks for a columnRank within group in for loop in RData transformation: from dyadic to observational data in RGetting map from purrr to work with paste0

                    Are all passive ability checks floors for active ability checks?Does passive perception supersede active perception?Which skills can be used passively?Active Opposition with Free-Form Professions in Fate5E Trap/Ambush/Stealth Mechanics VS Passive Perception ConfusionInteraction between perception and stealth in obscured conditionsHow does Keen Sight affect Passive Perception?Are all d20 rolls either attacks, saves or ability checks?Can players declare that they are making a specific ability check?Can I see a Hidden creature that is not obscured at all?Can a Stealth check ever be made passively?Is this alternate version of the Observant feat balanced?What is the minimum amount of skill points per HD?