You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 41747f4
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: docs/day3/pandas.rst
+53-14Lines changed: 53 additions & 14 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -236,10 +236,10 @@ To know if Pandas is the right tool for your job, you can consult the flowchart
236
236
237
237
We will also have a short session after this on plotting with Seaborn, a package for easily making publication-ready statistical plots with Pandas data structures.
238
238
239
-
Basic Data Types and Object Classes
240
-
-----------------------------------
239
+
Important Data Types and Object Classes
240
+
---------------------------------------
241
241
242
-
The main object classes of Pandas are ``Series`` and ``DataFrame``. There is also a separate object class called ``Index`` for the row indexes/labels and column labels, if applicable. Data that you load from file will mainly be loaded into either Series or DataFrames. Indexes are typically extracted later.
242
+
The main object classes of Pandas are ``Series`` and ``DataFrame``. There is also a separate object class called ``Index`` for the row indexes/labels and column labels, if applicable. Data that you load from file will mainly be loaded into either Series or DataFrames. Indexes are typically extracted later if needed.
243
243
244
244
* ``pandas.Series(data, index=None, name=None, ...)`` instantiates a 1D array with customizable indexes (labels) attached to every entry for easy access, and optionally a name for later addition to a DataFrame as a column.
245
245
@@ -252,7 +252,8 @@ The main object classes of Pandas are ``Series`` and ``DataFrame``. There is als
252
252
253
253
For the rest of this lesson, example DataFrames will be abbreviated as ``df`` in code snippets (and example Series, if they appear, will be abbreviated as ``ser``).
The API reference in the `official Pandas documentation <https://pandas.pydata.org/docs/user_guide/index.html>`_ shows hundreds of methods and attributes for Series and DataFrames. The following is a very brief list of the most important attributes and what they output.
258
259
@@ -279,27 +280,30 @@ There are also specialized datatypes for, e.g. saving on memory or performing wi
279
280
280
281
This is far from an exhaustive list.
281
282
282
-
Input/Output and Making DataFrames from Scratch
283
-
-----------------------------------------------
283
+
Loading or Creating DataFrames
284
+
------------------------------
284
285
285
-
Most of the time, Series and DataFrames will be loaded from files, not made from scratch. The following table lists I/O functions for a few of the most common data formats; `the full table with links to the documentation pages for each function can be found here. <https://pandas.pydata.org/pandas-docs/stable/user_guide/io.html>`__ Input and output functions are sometimes called readers and writers, respectively. The ``read_csv()`` is by far the most commonly used since it can read any text file with a specified delimiter (comma, tab, or otherwise).
286
+
Loading DataFrames from File
287
+
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
288
+
289
+
Most of the time, Series and DataFrames will be loaded from files, not made from scratch. To review, the following table lists I/O functions for a few of the most common data formats; `the full table with links to the documentation pages for each function can be found here. <https://pandas.pydata.org/pandas-docs/stable/user_guide/io.html>`__ Input and output functions are sometimes called readers and writers, respectively. The ``read_csv()`` is by far the most commonly used since it can read any text file with a specified delimiter (comma, tab, or otherwise).
This is far from a complete list, and most of these functions have several dozen possible kwargs. *Most kwargs in a given reader function also appear in the corresponding writer function, and serve the same purpose.* It is left to the reader to determine which kwargs are needed. As with NumPy's ``genfromtxt()`` function, most of the *text* readers above, and the excel reader, have kwargs that let you choose to load only some of the data.
302
+
Most of these functions have several dozen possible kwargs. It is left to the reader to determine which kwargs are needed. *Most kwargs in a given reader function also appear in the corresponding writer function, and serve the same purpose.* As with NumPy's ``genfromtxt()`` function, most of the *text* readers above, and the excel reader, have kwargs that let you choose to load only some of the data.
300
303
301
-
In the example below, a CSV file called "exoplanets_5250_EarthUnits.csv" in the current working directory is read into the DataFrame ``df`` and then written out to a plain text file where decimals are rendered with commas, the delimiter is the pipe character, and the indexes are preserved as the first column.
304
+
Most of the above formats were chosen not only because they are common, but because, apart from ``read_excel()``, these support **chunking** for data sets that are larger than memory.
302
305
306
+
In the example below, a CSV file called "exoplanets_5250_EarthUnits_fixed.csv" in the current working directory is read into the DataFrame ``df`` and then written out to a plain text file where decimals are rendered with commas, the delimiter is the pipe character, and the indexes are preserved as the first column.
303
307
304
308
.. challenge::
305
309
@@ -313,6 +317,9 @@ In the example below, a CSV file called "exoplanets_5250_EarthUnits.csv" in the
313
317
314
318
In most reader functions, including ``index_col=0`` sets the first column as the row labels, and the first row is assumed to contain the list of column names by default. If you forget to set one of the columns as the list of row indexes during import, you can do it later with ``df.set_index('column_name')``.
315
319
320
+
Creating DataFrames in Python
321
+
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
322
+
316
323
Building a DataFrame or Series from scratch is also easy. Lists and arrays can be converted directly to Series and DataFrames, respectively.
317
324
318
325
* Both ``pd.Series()`` and ``pd.DataFrame()`` have an ``index`` kwarg to assign a list of numbers, names, times, or other hashable keys to each row.
@@ -334,19 +341,16 @@ Building a DataFrame or Series from scratch is also easy. Lists and arrays can b
334
341
335
342
It is also possible to convert DataFrames and Series to NumPy arrays (with or without the indexes), dictionaries, record arrays, or strings with the methods ``.to_numpy()``, ``.to_dict()``, ``to_records()``, and ``to_string()``, respectively.
336
343
337
-
338
344
Inspection and Memory Usage
339
345
---------------------------
340
346
341
-
Review of Inspect
342
347
The main data inspection functions for DataFrames (and Series) are as follows:
343
348
344
349
* ``df.head()`` (or ``df.tail()``) prints first (or last) 5 rows of data with row and column labels, or accepts an integer argument to print a different number of rows.
345
350
* ``df.info()`` prints the number of rows with their first and last index values; titles, index numbers, valid data counts, and datatypes of columns; and the estimated size of ``df`` in memory. **Note:** do not rely on this memory estimate if your dataframe contains non-numeric data (see below).
346
351
* ``df.describe()`` prints summary statistics for all the numerical columns in ``df``.
347
352
* ``df.nunique()`` prints counts of the unique values in each column.
348
353
* ``df.value_counts()`` prints each unique value and the number of of occurrences for every combination of row and column values for as many of each as are selected (usually applied to just a couple of columns at a time at most)
349
-
* ``df.sample()`` randomly selects a given number of rows ``n=nrows``, or a decimal fraction ``frac`` of the total number of rows.
350
354
* ``df.memory_usage()`` returns the estimated memory usage per column (see important notes below).
351
355
352
356
.. important:: The ``memory_usage()`` Function
@@ -362,13 +366,48 @@ The main data inspection functions for DataFrames (and Series) are as follows:
.. admonition:: "Selection and Preprocessing Cheatsheet"
371
+
:collapsible:
372
+
373
+
Below is a table of the syntax for how to select or assign different subsets or cross-sections of a DataFrame. `The complete guide, including how to select data by conditions, can be found at this link. <https://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html>`__
The following table describes basic functions for finding, removing, and replacing missing or unwanted data, which will be necessary ahead of any machine learning applications. Pandas has its own functions for detecting missing data in order to detect both regular ``NaN``s and the datetime equivalent, ``NaT``. Any of the following functions will work on individual columns or any other subset of the DataFrame as well as the whole. `Click here for more information on handling missing or invalid data in Pandas. <https://pandas.pydata.org/pandas-docs/stable/user_guide/missing_data.html>`__
**Categorical data.** As the memory usage outputs show in the example above, a single 5-8-letter word uses almost 8 times as much memory as a 64-bit float. The ``Categorical`` datatype provides, among other benefits, a way to get the memory savings of a dummy variable array without having to create one, as long as the number of unique values is much smaller than the number of entries in the column(s) to be converted to ``Categorical`` type. Internally, the ``Categorical`` type maps all the unique values of a column to short numerical codes in the column's place in memory, stores the codes in the smallest integer format that fits the largest-valued code, and only converts the codes to the associated strings when the data are printed.
410
+
**Categorical data.** ``Categorical`` typemaps all the unique values of a column to short numerical codes in the column's place in memory, stores the codes in the smallest integer format that fits the largest-valued code, and only converts the codes to the associated strings when the data are printed. This data type is extremely efficient when the number of unique values are small relative to the size of the data set, but it is not recommended when half or more of the data values are unique.
372
411
373
412
* To convert a column in an existing Dataframe, simply set that column equal to itself with ``.astype('category')`` at the end. If defining a new Series that you want to be categorical, simply include ``dtype='category'``.
374
413
* To get attributes or call methods of ``Categorical`` data, use the ``.cat`` accessor followed by the attribute or method. E.g., to get the category names as an index object, use ``df['cat_col'].cat.categories``.
0 commit comments