'VfiY\ dZddlmZddlZddlZddlZddlmZmZm Z ddl Z ddl m Z m Z ddl mZddlmZddlmZdd lmZdd lmZdd lmZdd lmZdd lmZmZddlmZddl m!Z!ddl"m#Z#m$Z$m%Z%m&Z&m'Z'erddl(m)Z)m*Z*m+Z+m,Z,m-Z-dAdZ. dBdCd#Z/Gd$dZ0Gd%d&e0Z1Gd'd(e0Z2eed) dDdEd7Z3eed)d*ddej4ej4ddfdFd@Z5dS)Gz parquet compat ) annotationsN) TYPE_CHECKINGAnyLiteral)catch_warningsfilterwarnings) _get_option)lib)import_optional_dependencyAbstractMethodError)doc)find_stack_level)check_dtype_backend) DataFrame get_option) _shared_docs)arrow_table_to_pandas) IOHandles get_handle is_fsspec_urlis_urlstringify_path) DtypeBackendFilePath ReadBufferStorageOptions WriteBufferenginestrreturnBaseImplcd|dkrtd}|dkr_ttg}d}|D]:} |cS#t$r}|dt |zz }Yd}~3d}~wwxYwtd||dkrtS|dkrtSt d ) zreturn our implementationautozio.parquet.enginez - NzUnable to find a usable engine; tried using: 'pyarrow', 'fastparquet'. A suitable version of pyarrow or fastparquet is required for parquet support. Trying to import the above resulted in these errors:pyarrow fastparquetz.engine must be one of 'pyarrow', 'fastparquet')r PyArrowImplFastParquetImpl ImportErrorr ValueError)rengine_classes error_msgs engine_classerrs @c:\PYTHON\DbComparer\venv\Lib\site-packages\pandas/io/parquet.py get_enginer14s /00 %7 * 1 1L 1#|~~%%% 1 1 1gC00  1        }} =    E F FFs = A&A!!A&rbFpath1FilePath | ReadBuffer[bytes] | WriteBuffer[bytes]fsrstorage_optionsStorageOptions | Nonemodeis_dirboolVtuple[FilePath | ReadBuffer[bytes] | WriteBuffer[bytes], IOHandles[bytes] | None, Any]c`t|}|tdd}tdd}|'t||jr|rt dnA|t||jjrn$tdt|j t|r||Ttd}td} |j |\}}n#t|j f$rYnwxYw|'td}|jj|fi|pi\}}n&|r$t!|r|d krtd d} |sR|sPt|t"r;t$j|st+||d | } d}| j}|| |fS) zFile handling for PyArrow.Nz pyarrow.fsignore)errorsfsspecz8storage_options not supported with a pyarrow FileSystem.z9filesystem must be a pyarrow or fsspec FileSystem, not a r&r2z8storage_options passed with buffer, or non-supported URLFis_textr6)rr isinstance FileSystemNotImplementedErrorspecAbstractFileSystemr+type__name__rfrom_uri TypeError ArrowInvalidcore url_to_fsrr osr3isdirrhandle) r3r5r6r8r9path_or_handlepa_fsr?pahandless r0_get_path_or_handlerUVs2$D))N ~*<III+HXFFF  B0@!A!A  )N  Jr6;3Q$R$R  -b*-- ^$$U  "+I66B.|<%>t%D%D"NNr/     :/99F!6!6""#2#8b"" B U&"8"8UDDLLSTTTG  ( ( ~s + + ( n-- ( D%     7B &&sC..DDc8eZdZed dZd dZd d dZdS) r"dfrr!NonecNt|tstddS)Nz+to_parquet only supports IO with DataFrames)rBrr+)rWs r0validate_dataframezBaseImpl.validate_dataframes0"i(( LJKK K L Lc  t|Nr )selfrWr3 compressionkwargss r0writezBaseImpl.write!$'''r[Nc  t|r]r )r^r3columnsr`s r0readz BaseImpl.readrbr[)rWrr!rX)rWrr])r!r)rH __module__ __qualname__ staticmethodrZrarer[r0r"r"scLLL\L(((((((((((r[cJeZdZddZ dddZdddejddfddZdS)r(r!rXcFtddddl}ddl}||_dS)Nr&z(pyarrow is required for parquet support.extrar)r pyarrow.parquet(pandas.core.arrays.arrow.extension_typesapi)r^r&pandass r0__init__zPyArrowImpl.__init__sF" G      8777r[snappyNrWrr3FilePath | WriteBuffer[bytes]r_ str | Noneindex bool | Noner6r7partition_colslist[str] | Nonec R||d|ddi} ||| d<|jjj|fi| } |jrBdt j|ji} | jj } i| | } | | } t|||d|du\}}}t|tjrlt|dr\t|jt"t$fr;t|jt$r|j}n|j} ||jjj| |f|||d|n|jjj| |f||d|||dSdS#||wwxYw) Nschemapreserve_index PANDAS_ATTRSwb)r6r8r9name)r_rx filesystem)r_r)rZpoprpTable from_pandasattrsjsondumpsr{metadatareplace_schema_metadatarUrBioBufferedWriterhasattrrr bytesdecodeparquetwrite_to_dataset write_tableclose)r^rWr3r_rvr6rxrr`from_pandas_kwargstable df_metadataexisting_metadatamerged_metadatarQrTs r0razPyArrowImpl.writes( ###.6 8T8R8R-S  38 / 0**2DD1CDD 8 C)4:bh+?+?@K % 5 B!2BkBO11/BBE.A  +!- / / / + ~r'8 9 9 5// 5>.e == 5 .-u55 5!/!4!;!;!=!=!/!4 )1 1"!,#1)   - ,"!,)   " #"w" #s 7>>"*/":?"KK#':k#:#:FL" #w" #s0*D!&)B D!BD!"B#A&D!!D:r!rXrsNNNN)rWrr3rtr_rurvrwr6r7rxryr!rX)rr:rrr6r7r!r)rHrfrgrrrar no_defaultrerir[r0r(r(s    #+!15+/@ @ @ @ @ J$)69n157 7 7 7 7 7 7 r[r(c<eZdZddZ ddd Z ddd ZdS)r)r!rXc6tdd}||_dS)Nr'z,fastparquet is required for parquet support.rl)r rp)r^r's r0rrzFastParquetImpl.__init__+s+1 !O   r[rsNrWrr_*Literal['snappy', 'gzip', 'brotli'] | Noner6r7c  ||d|vr|tdd|vr|d}|d|d<|tdt |}t |rt d fd|d<nrtd td 5|jj ||f|||d |ddddS#1swxYwYdS) N partition_onzYCannot use both partition_on and partition_cols. Use partition_cols for partitioning datahive file_scheme9filesystem is not implemented for the fastparquet engine.r?cJj|dfipiS)Nr~)open)r3_r?r6s r0z'FastParquetImpl.write..Vs8+&+d33.4"33dffr[ open_withz?storage_options passed with file object or non-fsspec file pathT)record)r_ write_indexr) rZr+rrDrrr rrpra) r^rWr3r_rvrxr6rr`r?s ` @r0razFastParquetImpl.write3s ### V # #(BK  V # ##ZZ77N  %$*F= !  !%K  d##    /99F#####F;   Q 4 ( ( (   DHN (!+                          s6CC #C c i}|dd}|dtj} d|d<|rtd| tjurtd|t dt |}d} t |r)td} | j|d fi|pij |d <nNt|tr9tj |st|d d| } | j} |jj|fi|} | jd ||d || | SS#| | wwxYw)NrFr pandas_nullszNThe 'use_nullable_dtypes' argument is not supported for the fastparquet enginezHThe 'dtype_backend' argument is not supported for the fastparquet enginerr?r2r5r@)rdrri)rr rr+rDrrr rr5rBr rNr3rOrrPrp ParquetFile to_pandasr) r^r3rdrr6rr`parquet_kwargsrrrTr? parquet_files r0rezFastParquetImpl.readhs*,$jj)>FF ?CNCC ).~&  %   . .%   !%K d##    "/99F#.6;tT#U#Uo>SQS#U#U#XN4 c " " "27==+>+> "!dE?G>D /48/GGGGL)<)U'7UUfUU" #w" #s "EE(rr)rWrr_rr6r7r!rX)NNNN)r6r7r!r)rHrfrgrrrarerir[r0r)r)*s|CK1533333p15 0 0 0 0 0 0 0 r[r))r6r$rsrWr$FilePath | WriteBuffer[bytes] | Noner_rurvrwrxryr bytes | Nonec t|tr|g}t|} |tjn|} | j|| f|||||d||0t| tjsJ| SdS)a Write a DataFrame to the parquet format. Parameters ---------- df : DataFrame path : str, path object, file-like object, or None, default None String, path object (implementing ``os.PathLike[str]``), or file-like object implementing a binary ``write()`` function. If None, the result is returned as bytes. If a string, it will be used as Root Directory path when writing a partitioned dataset. The engine fastparquet does not accept file-like objects. engine : {{'auto', 'pyarrow', 'fastparquet'}}, default 'auto' Parquet library to use. If 'auto', then the option ``io.parquet.engine`` is used. The default ``io.parquet.engine`` behavior is to try 'pyarrow', falling back to 'fastparquet' if 'pyarrow' is unavailable. When using the ``'pyarrow'`` engine and no storage options are provided and a filesystem is implemented by both ``pyarrow.fs`` and ``fsspec`` (e.g. "s3://"), then the ``pyarrow.fs`` filesystem is attempted first. Use the filesystem keyword with an instantiated fsspec filesystem if you wish to use its implementation. compression : {{'snappy', 'gzip', 'brotli', 'lz4', 'zstd', None}}, default 'snappy'. Name of the compression to use. Use ``None`` for no compression. index : bool, default None If ``True``, include the dataframe's index(es) in the file output. If ``False``, they will not be written to the file. If ``None``, similar to ``True`` the dataframe's index(es) will be saved. However, instead of being saved as values, the RangeIndex will be stored as a range in the metadata so it doesn't require much space and is faster. Other indexes will be included as columns in the file output. partition_cols : str or list, optional, default None Column names by which to partition the dataset. Columns are partitioned in the order they are given. Must be None if path is not a string. {storage_options} filesystem : fsspec or pyarrow filesystem, default None Filesystem object to use when reading the parquet file. Only implemented for ``engine="pyarrow"``. .. versionadded:: 2.1.0 kwargs Additional keyword arguments passed to the engine Returns ------- bytes if no path argument is provided else None N)r_rvrxr6r)rBr r1rBytesIOragetvalue) rWr3rr_rvr6rxrr`impl path_or_bufs r0 to_parquetrsB.#&&*() f  DAESWKDJ   %'       |+rz22222##%%%tr[FilePath | ReadBuffer[bytes]rdrbool | lib.NoDefaultrrr&list[tuple] | list[list[tuple]] | Nonec t|} |tjur4d} |dur| dz } tj| t t nd}t|| j|f||||||d|S)a Load a parquet object from the file path, returning a DataFrame. Parameters ---------- path : str, path object or file-like object String, path object (implementing ``os.PathLike[str]``), or file-like object implementing a binary ``read()`` function. The string could be a URL. Valid URL schemes include http, ftp, s3, gs, and file. For file URLs, a host is expected. A local file could be: ``file://localhost/path/to/table.parquet``. A file URL can also be a path to a directory that contains multiple partitioned parquet files. Both pyarrow and fastparquet support paths to directories as well as file URLs. A directory path could be: ``file://localhost/path/to/tables`` or ``s3://bucket/partition_dir``. engine : {{'auto', 'pyarrow', 'fastparquet'}}, default 'auto' Parquet library to use. If 'auto', then the option ``io.parquet.engine`` is used. The default ``io.parquet.engine`` behavior is to try 'pyarrow', falling back to 'fastparquet' if 'pyarrow' is unavailable. When using the ``'pyarrow'`` engine and no storage options are provided and a filesystem is implemented by both ``pyarrow.fs`` and ``fsspec`` (e.g. "s3://"), then the ``pyarrow.fs`` filesystem is attempted first. Use the filesystem keyword with an instantiated fsspec filesystem if you wish to use its implementation. columns : list, default=None If not None, only these columns will be read from the file. {storage_options} .. versionadded:: 1.3.0 use_nullable_dtypes : bool, default False If True, use dtypes that use ``pd.NA`` as missing value indicator for the resulting DataFrame. (only applicable for the ``pyarrow`` engine) As new dtypes are added that support ``pd.NA`` in the future, the output with this option will change to use those dtypes. Note: this is an experimental option, and behaviour (e.g. additional support dtypes) may change without notice. .. deprecated:: 2.0 dtype_backend : {{'numpy_nullable', 'pyarrow'}}, default 'numpy_nullable' Back-end data type applied to the resultant :class:`DataFrame` (still experimental). Behaviour is as follows: * ``"numpy_nullable"``: returns nullable-dtype-backed :class:`DataFrame` (default). * ``"pyarrow"``: returns pyarrow-backed nullable :class:`ArrowDtype` DataFrame. .. versionadded:: 2.0 filesystem : fsspec or pyarrow filesystem, default None Filesystem object to use when reading the parquet file. Only implemented for ``engine="pyarrow"``. .. versionadded:: 2.1.0 filters : List[Tuple] or List[List[Tuple]], default None To filter out data. Filter syntax: [[(column, op, val), ...],...] where op is [==, =, >, >=, <, <=, !=, in, not in] The innermost tuples are transposed into a set of filters applied through an `AND` operation. The outer list combines these sets of filters through an `OR` operation. A single list of tuples can also be used, meaning that no `OR` operation between set of filters is to be conducted. Using this argument will NOT result in row-wise filtering of the final partitions unless ``engine="pyarrow"`` is also specified. For other engines, filtering is only performed at the partition level, that is, to prevent the loading of some row-groups and/or files. .. versionadded:: 2.1.0 **kwargs Any additional kwargs are passed to the engine. Returns ------- DataFrame See Also -------- DataFrame.to_parquet : Create a parquet object that serializes a DataFrame. Examples -------- >>> original_df = pd.DataFrame( ... {{"foo": range(5), "bar": range(5, 10)}} ... ) >>> original_df foo bar 0 0 5 1 1 6 2 2 7 3 3 8 4 4 9 >>> df_parquet_bytes = original_df.to_parquet() >>> from io import BytesIO >>> restored_df = pd.read_parquet(BytesIO(df_parquet_bytes)) >>> restored_df foo bar 0 0 5 1 1 6 2 2 7 3 3 8 4 4 9 >>> restored_df.equals(original_df) True >>> restored_bar = pd.read_parquet(BytesIO(df_parquet_bytes), columns=["bar"]) >>> restored_bar bar 0 5 1 6 2 7 3 8 4 9 >>> restored_bar.equals(original_df[['bar']]) True The function uses `kwargs` that are passed directly to the engine. In the following example, we use the `filters` argument of the pyarrow engine to filter the rows of the DataFrame. Since `pyarrow` is the default engine, we can omit the `engine` argument. Note that the `filters` argument is implemented by the `pyarrow` engine, which can benefit from multithreading and also potentially be more economical in terms of memory. >>> sel = [("foo", ">", 2)] >>> restored_part = pd.read_parquet(BytesIO(df_parquet_bytes), filters=sel) >>> restored_part foo bar 0 3 8 1 4 9 zYThe argument 'use_nullable_dtypes' is deprecated and will be removed in a future version.TzFUse dtype_backend='numpy_nullable' instead of use_nullable_dtype=True.) stacklevelF)rdrr6rrr) r1r rwarningswarn FutureWarningrrre) r3rrdr6rrrrr`rmsgs r0 read_parquetrsr f  D#.00 #  $ & & X C  c=5E5G5GHHHHH# &&& 49  '/#      r[)rr r!r")Nr2F) r3r4r5rr6r7r8r r9r:r!r;)Nr$rsNNNN)rWrr3rrr r_rurvrwr6r7rxryrrr!r)r3rrr rdryr6r7rrrrrrrrr!r)6__doc__ __future__rrrrNtypingrrrrrrpandas._config.configr pandas._libsr pandas.compat._optionalr pandas.errorsr pandas.util._decoratorsrpandas.util._exceptionsrpandas.util._validatorsrrqrrpandas.core.shared_docsrpandas.io._utilrpandas.io.commonrrrrrpandas._typingrrrrrr1rUr"r(r)rrrrir[r0rs""""""   .----->>>>>>------''''''444444777777100000111111GGGGJ.2 <'<'<'<'<'~ ( ( ( ( ( ( ( (E E E E E (E E E Pn n n n n hn n n b\"3455526&-1'+UUUU65Up\"34555 $-10325.6:qqqq65qqqr[