B ª`R…ã @sDdZdZddlZddlmZddlZddlZddlZdZyddl Z dd„Z WnFe k r’yddl Z dd„Z Wne k rŒdd„Z YnXYnXy ddl Z Wne k r´YnXd Zd ZeƒZe e d ¡ej¡e e d ¡ej¡d œee<e eej¡e eej¡d œee<Gd d„deƒZGdd„dƒZGdd„dƒZdS)aBBeautiful Soup bonus library: Unicode, Dammit This library converts a bytestream to Unicode through any means necessary. It is heavily based on code from Mark Pilgrim's Universal Feed Parser. It works best on XML and HTML, but it does not rewrite the XML or HTML to reflect a new encoding; that's the tree builder's job. ÚMITéN)Úcodepoint2namecCst|tƒrdSt |¡dS)NÚencoding)Ú isinstanceÚstrÚcchardetÚdetect)Ús©r új/private/var/folders/fw/jsxvvqfs4sz4tdnfdvg5typ5vk77qg/T/pip-install-p7nfy4dm/beautifulsoup4/bs4/dammit.pyÚchardet_dammits r cCst|tƒrdSt |¡dS)Nr)rrÚchardetr)r r r r r "s cCsdS)Nr )r r r r r *sz$^\s*<\?.*encoding=['"](.*?)['"].*\?>z0<\s*meta[^>]+charset\s*=\s*["']?([^>]*?)[ /;'">]Úascii)ÚhtmlÚxmlc@s”eZdZdZdd„Zeƒ\ZZZdddddd œZe   d ¡Z e   d ¡Z e d d „ƒZe dd„ƒZe dd„ƒZe ddd„ƒZe ddd„ƒZe dd„ƒZdS)ÚEntitySubstitutionzFThe ability to substitute XML or HTML entities for certain characters.cCsxi}i}g}dg}xFtt ¡ƒ|D]2\}}t|ƒ}|dkrN| |¡|||<|||<q$Wdd |¡}||t |¡fS)N)é'Úapos)é"rz[%s]Ú)ÚlistrÚitemsÚchrÚappendÚjoinÚreÚcompile)ÚlookupZreverse_lookupZcharacters_for_reÚextraÚ codepointÚnameÚ characterZ re_definitionr r r Ú_populate_class_variablesGs  z,EntitySubstitution._populate_class_variablesrÚquotÚampÚltÚgt)ú'ú"ú&ú<ú>z&([<>]|&(?!#\d+;|#x[0-9a-fA-F]+;|\w+;))z([<>&])cCs|j | d¡¡}d|S)ziUsed with a regular expression to substitute the appropriate HTML entity for a special character.rz&%s;)ÚCHARACTER_TO_HTML_ENTITYÚgetÚgroup)ÚclsÚmatchobjÚentityr r r Ú_substitute_html_entityqsz*EntitySubstitution._substitute_html_entitycCs|j| d¡}d|S)zhUsed with a regular expression to substitute the appropriate XML entity for a special character.rz&%s;)ÚCHARACTER_TO_XML_ENTITYr.)r/r0r1r r r Ú_substitute_xml_entityxsz)EntitySubstitution._substitute_xml_entitycCs6d}d|kr*d|kr&d}| d|¡}nd}|||S)a*Make a value into a quoted XML attribute, possibly escaping it. Most strings will be quoted using double quotes. Bob's Bar -> "Bob's Bar" If a string contains double quotes, it will be quoted using single quotes. Welcome to "my bar" -> 'Welcome to "my bar"' If a string contains both single and double quotes, the double quotes will be escaped, and the string will be quoted using double quotes. Welcome to "Bob's Bar" -> "Welcome to "Bob's bar" r(r'z")Úreplace)ÚselfÚvalueZ quote_withZ replace_withr r r Úquoted_attribute_valuesz)EntitySubstitution.quoted_attribute_valueFcCs"|j |j|¡}|r| |¡}|S)a Substitute XML entities for special XML characters. :param value: A string to be substituted. The less-than sign will become <, the greater-than sign will become >, and any ampersands will become &. If you want ampersands that appear to be part of an entity definition to be left alone, use substitute_xml_containing_entities() instead. :param make_quoted_attribute: If True, then the string will be quoted, as befits an attribute value. )ÚAMPERSAND_OR_BRACKETÚsubr4r8)r/r7Úmake_quoted_attributer r r Úsubstitute_xml¤s   z!EntitySubstitution.substitute_xmlcCs"|j |j|¡}|r| |¡}|S)aŸSubstitute XML entities for special XML characters. :param value: A string to be substituted. The less-than sign will become <, the greater-than sign will become >, and any ampersands that are not part of an entity defition will become &. :param make_quoted_attribute: If True, then the string will be quoted, as befits an attribute value. )ÚBARE_AMPERSAND_OR_BRACKETr:r4r8)r/r7r;r r r Ú"substitute_xml_containing_entities¹s   z5EntitySubstitution.substitute_xml_containing_entitiescCs|j |j|¡S)aReplace certain Unicode characters with named HTML entities. This differs from data.encode(encoding, 'xmlcharrefreplace') in that the goal is to make the result more readable (to those with ASCII displays) rather than to recover from errors. There's absolutely nothing wrong with a UTF-8 string containg a LATIN SMALL LETTER E WITH ACUTE, but replacing that character with "é" will make it more readable to some people. :param s: A Unicode string. )ÚCHARACTER_TO_HTML_ENTITY_REr:r2)r/r r r r Úsubstitute_htmlÏsz"EntitySubstitution.substitute_htmlN)F)F)Ú__name__Ú __module__Ú __qualname__Ú__doc__r"r,ZHTML_ENTITY_TO_CHARACTERr?r3rrr=r9Ú classmethodr2r4r8r<r>r@r r r r rDs$      %  rc@sHeZdZdZddd„Zdd„Zedd „ƒZed d „ƒZ edd d „ƒZ dS)ÚEncodingDetectora^Suggests a number of possible encodings for a bytestring. Order of precedence: 1. Encodings you specifically tell EncodingDetector to try first (the override_encodings argument to the constructor). 2. An encoding declared within the bytestring itself, either in an XML declaration (if the bytestring is to be interpreted as an XML document), or in a tag (if the bytestring is to be interpreted as an HTML document.) 3. An encoding detected through textual analysis by chardet, cchardet, or a similar external library. 4. UTF-8. 5. Windows-1252. NFcCsN|pg|_|pg}tdd„|Dƒƒ|_d|_||_d|_| |¡\|_|_dS)a€Constructor. :param markup: Some markup in an unknown encoding. :param override_encodings: These encodings will be tried first. :param is_html: If True, this markup is considered to be HTML. Otherwise it's assumed to be XML. :param exclude_encodings: These encodings will not be tried, even if they otherwise would be. cSsg|] }| ¡‘qSr )Úlower)Ú.0Úxr r r ú sz-EncodingDetector.__init__..N) Úoverride_encodingsÚsetÚexclude_encodingsÚchardet_encodingÚis_htmlÚdeclared_encodingÚstrip_byte_order_markÚmarkupÚsniffed_encoding)r6rRrKrOrMr r r Ú__init__õs zEncodingDetector.__init__cCs8|dk r4| ¡}||jkrdS||kr4| |¡dSdS)zÕShould we even bother to try this encoding? :param encoding: Name of an encoding. :param tried: Encodings that have already been tried. This will be modified as a side effect. NFT)rGrMÚadd)r6rÚtriedr r r Ú_usable s  zEncodingDetector._usableccsÀtƒ}x |jD]}| ||¡r|VqW| |j|¡r>|jV|jdkrZ| |j|j¡|_| |j|¡rp|jV|jdkr†t |jƒ|_| |j|¡rœ|jVxdD]}| ||¡r¢|Vq¢WdS)zmYield a number of encodings that might work for this markup. :yield: A sequence of strings. N)zutf-8z windows-1252) rLrKrWrSrPÚfind_declared_encodingrRrOrNr )r6rVÚer r r Ú encodingss$        zEncodingDetector.encodingscCsþd}t|tƒr||fSt|ƒdkrT|dd…dkrT|dd…dkrTd}|dd…}n¢t|ƒdkr’|dd…dkr’|dd…dkr’d}|dd…}nd|dd …d kr´d }|d d…}nB|dd…d krÖd }|dd…}n |dd…dkröd}|dd…}||fS)z¶If a byte-order mark is present, strip it and return the encoding it implies. :param data: Some markup. :return: A 2-tuple (modified data, implied encoding) Néésþÿzzutf-16besÿþzutf-16leészutf-8sþÿzutf-32besÿþzutf-32le)rrÚlen)r/Údatarr r r rQ>s*  z&EncodingDetector.strip_byte_order_markc Csº|rt|ƒ}}nd}tdtt|ƒdƒƒ}t|tƒr@tt}ntt}|d}|d}d} |j||d�} | s€|r€|j||d�} | dk r”|  ¡d} | r¶t| tƒr®|   d d ¡} |   ¡SdS) a§Given a document, tries to find its declared encoding. An XML encoding is declared at the beginning of the document. An HTML encoding is declared in a tag, hopefully near the beginning of the document. :param markup: Some markup. :param is_html: If True, this markup is considered to be HTML. Otherwise it's assumed to be XML. :param search_entire_document: Since an encoding is supposed to declared near the beginning of the document, most of the time it's only necessary to search a few kilobytes of data. Set this to True to force this method to search the entire document. iigš™™™™™©?rrN)Úendposrrr5) r^ÚmaxÚintrÚbytesÚ encoding_resrÚsearchÚgroupsÚdecoderG) r/rRrOZsearch_entire_documentZ xml_endposZ html_endposÚresÚxml_reZhtml_rerPZdeclared_encoding_matchr r r rX\s(     z'EncodingDetector.find_declared_encoding)NFN)FF) rArBrCrDrTrWÚpropertyrZrErQrXr r r r rFás  $ rFc�@sæeZdZdZdddœZdddgZgdd gfd d „Zd d „Zdþdd„Zdÿdd„Z e dd„ƒZ dd„Z dd„Z dddddddd d!d"d#d$d%d&d'd&d&d(d)d*d+d,d-d.d/d0d1d2d3d&d4d5d6œ Zd7dd8d9d:d;dd?d@dAdBd&dCd&d&dDdDdEdEdFdGdHdIdJdKdLdMd&dNdOddPdQdRdSdTdUd@dVdWdXdYdPddZdGd[d\d]d^d_d`dadFd8dbdXdcdddedfd&dgdgdgdgdgdgdhdidjdjdjdjdkdkdkdkdldmdndndndndndFdndododododOdpdqdrdrdrdrdrdrdsdQdtdtdtdtdudududud[dvd[d[d[d[d[dwd[d`d`d`d`dxdpdxdyœ€Zdzd{d|d}d~dd€d�d‚dƒd„d…d†d‡dˆd‰dŠd‹dŒd�dŽd�d�d‘d’d“d”d•d–d—d˜d™dšd›dœd�dždŸd d¡d¢d£d¤d¥d¦d§d¨d©dªd«d¬d­d®d¯d°d±d²d³d´dµd¶d·d¸d¹dºd»d¼d½d¾d¿dÀdÁdÂdÃdÄdÅdÆdÇdÈdÉdÊdËdÌdÍdÎdÏdÐdÑdÒdÓdÔdÕdÖd×dØdÙdÚdÛdÜdÝdÞdßdàdádâdãdädådædçdèdédêdëdìdídîdïdðdñdòdódôœzZdõdöd÷gZedødøZedùdúZe�ddüdý„ƒZdS(Ú UnicodeDammitzÏA class for detecting the encoding of a *ML document and converting it to a Unicode string. If the source encoding is windows-1252, can replace MS smart quotes with their HTML or XML equivalents.z mac-romanz shift-jis)Ú macintoshzx-sjisú windows-1252z iso-8859-1z iso-8859-2NFcCsö||_g|_d|_||_t t¡|_t||||ƒ|_ t |t ƒsF|dkr`||_ t |ƒ|_ d|_dS|j j |_ d}x,|j jD] }|j j }| |¡}|dk rxPqxW|sâx@|j jD]4}|dkrÂ| |d¡}|dk rª|j d¡d|_PqªW||_ |sòd|_dS)aPConstructor. :param markup: A bytestring representing markup in an unknown encoding. :param override_encodings: These encodings will be tried first, before any sniffing code is run. :param smart_quotes_to: By default, Microsoft smart quotes will, like all other characters, be converted to Unicode characters. Setting this to 'ascii' will convert them to ASCII quotes instead. Setting it to 'xml' will convert them to XML entity references, and setting it to 'html' will convert them to HTML entity references. :param is_html: If True, this markup is considered to be HTML. Otherwise it's assumed to be XML. :param exclude_encodings: These encodings will not be considered, even if the sniffing code thinks they might make sense. FrNrr5zSSome characters could not be decoded, and were replaced with REPLACEMENT CHARACTER.T)Úsmart_quotes_toÚtried_encodingsZcontains_replacement_charactersrOÚloggingÚ getLoggerrAÚlogrFÚdetectorrrrRZunicode_markupÚoriginal_encodingrZÚ _convert_fromÚwarning)r6rRrKrnrOrMÚurr r r rT˜s>     zUnicodeDammit.__init__cCs�| d¡}|jdkr&|j |¡ ¡}nf|j |¡}t|ƒtkr„|jdkrfd ¡|d ¡d ¡}qŒd ¡|d ¡d ¡}n| ¡}|S)z[Changes a MS smart quote character to an XML or HTML entity, or an ASCII character.érrz&#xú;r)r)r.rnÚMS_CHARS_TO_ASCIIr-ÚencodeÚMS_CHARSÚtypeÚtuple)r6ÚmatchÚorigr:r r r Ú _sub_ms_charÙs     zUnicodeDammit._sub_ms_charÚstrictc Cs®| |¡}|r||f|jkr dS|j ||f¡|j}|jdk rf||jkrfd}t |¡}| |j |¡}y|  |||¡}||_||_ Wn"t k r¦}zdSd}~XYnX|jS)z|Attempt to convert the markup to the proposed encoding. :param proposed: The name of a character encoding. Ns([€-Ÿ])) Ú find_codecrorrRrnÚENCODINGS_WITH_SMART_QUOTESrrr:r�Ú _to_unicodertÚ Exception)r6ZproposedÚerrorsrRZsmart_quotes_reZsmart_quotes_compiledrwrYr r r ruês"     zUnicodeDammit._convert_fromcCs t|||ƒS)z}Given a string and its encoding, decodes the string into Unicode. :param encoding: The name of an encoding. )r)r6r_rr‡r r r r… szUnicodeDammit._to_unicodecCs|js dS|jjS)zhIf the markup is an HTML document, returns the encoding declared _within_ the document. N)rOrsrP)r6r r r Údeclared_html_encodingsz$UnicodeDammit.declared_html_encodingcCs`| |j ||¡¡pN|r*| | dd¡¡pN|r@| | dd¡¡pN|rL| ¡pN|}|r\| ¡SdS)z™Convert the name of a character set to a codec name. :param charset: The name of a character set. :return: The name of a codec. ú-rÚ_N)Ú_codecÚCHARSET_ALIASESr-r5rG)r6Úcharsetr7r r r rƒs zUnicodeDammit.find_codecc Cs<|s|Sd}yt |¡|}Wnttfk r6YnX|S)N)ÚcodecsrÚ LookupErrorÚ ValueError)r6r�Úcodecr r r r‹)s zUnicodeDammit._codec)ÚeuroZ20ACú )ÚsbquoZ201A)ÚfnofZ192)ÚbdquoZ201E)ÚhellipZ2026)ÚdaggerZ2020)ÚDaggerZ2021)ÚcircZ2C6)ÚpermilZ2030)ÚScaronZ160)ÚlsaquoZ2039)ÚOEligZ152ú?)z#x17DZ17D)ÚlsquoZ2018)ÚrsquoZ2019)ÚldquoZ201C)ÚrdquoZ201D)ÚbullZ2022)ÚndashZ2013)ÚmdashZ2014)ÚtildeZ2DC)ÚtradeZ2122)ÚscaronZ161)ÚrsaquoZ203A)ÚoeligZ153)z#x17EZ17E)ÚYumlr) ó€ó�ó‚óƒó„ó…ó†ó‡óˆó‰óŠó‹óŒó�óŽó�ó�ó‘ó’ó“ó”ó•ó–ó—ó˜ó™óšó›óœó�óžóŸZEURú,Úfz,,z...ú+z++ú^ú%ÚSr*ZOEÚZr'r(Ú*r‰z--ú~z(TM)r r+ZoeÚzÚYú!ÚcZGBPú$ZYENú|z..rz(th)z<>z1/4z1/2z3/4ÚAZAEÚCÚEÚIÚDÚNÚOÚUÚbÚBÚaZaerYÚiÚnú/Úy)€r­r®r¯r°r±r²r³r´rµr¶r·r¸r¹rºr»r¼r½r¾r¿rÀrÁrÂrÃrÄrÅrÆrÇrÈrÉrÊrËrÌó ó¡ó¢ó£ó¤ó¥ó¦ó§ó¨ó©óªó«ó¬ó­ó®ó¯ó°ó±ó²ó³ó´óµó¶ó·ó¸ó¹óºó»ó¼ó½ó¾ó¿óÀóÁóÂóÃóÄóÅóÆóÇóÈóÉóÊóËóÌóÍóÎóÏóÐóÑóÒóÓóÔóÕóÖó×óØóÙóÚóÛóÜóÝóÞóßóàóáóâóãóäóåóæóçóèóéóêóëóìóíóîóïóðóñóòóóóôóõóöó÷óøóùóúóûóüóýóþóÿs€s‚sÆ’s„s…s†s‡sˆs‰sÅ s‹sÅ’sŽs‘s’s“sâ€�s•s–s—sËœsâ„¢sÅ¡s›sÅ“sžsŸs s¡s¢s£s¤sÂ¥s¦s§s¨s©sªs«s¬s­s®s¯s°s±s²s³s´sµs¶s·s¸s¹sºs»s¼s½s¾s¿sÀsÃ�sÂsÃsÄsÃ…sÆsÇsÈsÉsÊsËsÃŒsÃ�sÃŽsÃ�sÃ�sÑsÃ’sÓsÔsÕsÖs×sØsÙsÚsÛsÜsÃ�sÞsßsàròsâsãsäsÃ¥sæsçsèsésêsësìsísîsïsðsñsòsósôsõsös÷søsùsúsûsüsýsþ)zé€é‚éƒé„é…é†é‡éˆé‰éŠé‹éŒéŽé‘é’é“é”é•é–é—é˜é™éšé›éœéžéŸé é¡é¢é£é¤é¥é¦é§é¨é©éªé«é¬é­é®é¯é°é±é²é³é´éµé¶é·é¸é¹éºé»é¼é½é¾é¿éÀéÁéÂéÃéÄéÅéÆéÇéÈéÉéÊéËéÌéÍéÎéÏéÐéÑéÒéÓéÔéÕéÖé×éØéÙéÚéÛéÜéÝéÞéßéàéáéâéãéäéåéæéçéèéééêéëéìéíéîéïéðéñéòéóéôéõéöé÷éøéùéúéûéüéýéþ)rŽr«r\)r¬r»r])r¼rÀr[réÿÿÿÿrxÚutf8c Cs"| dd¡ ¡dkrtdƒ‚| ¡dkr0tdƒ‚g}d}d}xº|t|ƒkrö||}t|tƒsdt|ƒ}||jkrª||jkrªxz|j D]$\}} } ||kr€|| kr€|| 7}Pq€Wq>|dkrì||j krì|  |||…¡|  |j |¡|d 7}|}q>|d 7}q>W|dk�r|S|  ||d …¡d   |¡S) aFix characters from one encoding embedded in some other encoding. Currently the only situation supported is Windows-1252 (or its subset ISO-8859-1), embedded in UTF-8. :param in_bytes: A bytestring that you suspect contains characters from multiple encodings. Note that this _must_ be a bytestring. If you've already converted the document to Unicode, you're too late. :param main_encoding: The primary encoding of `in_bytes`. :param embedded_encoding: The encoding that was used to embed characters in the main document. :return: A bytestring in which `embedded_encoding` characters have been converted to their `main_encoding` equivalents. rŠr‰)z windows-1252Ú windows_1252zPWindows-1252 and ISO-8859-1 are the only currently supported embedded encodings.)rÌzutf-8z4UTF-8 is the only currently supported main encoding.rrQrxNó) r5rGÚNotImplementedErrorr^rrbÚordÚFIRST_MULTIBYTE_MARKERÚLAST_MULTIBYTE_MARKERÚMULTIBYTE_MARKERS_AND_SIZESÚWINDOWS_1252_TO_UTF8rr) r/Zin_bytesZ main_encodingZembedded_encodingZ byte_chunksZ chunk_startÚposÚbyteÚstartÚendÚsizer r r Ú detwingleis<      zUnicodeDammit.detwingle)r‚)r‚)rÌrm)rArBrCrDrŒr„rTr�rur…rjrˆrƒr‹r|rzrÔrÓrÑrÒrErÚr r r r rk…s`@       rk)rDÚ __license__rŽÚ html.entitiesrrrpÚstringZ chardet_typerr Ú ImportErrorr Z iconv_codecZ xml_encodingZ html_metaÚdictrdrr{rårcrÚobjectrrFrkr r r r Ús@     %