str.replace, es sei denn, auf string folgt ein bestimmter Text
Oct 25 2020
Überarbeitung der vorherigen Frage:
Wie kann ich alle "," (dh Komma dann Leerzeichen) durch "_" ersetzen, außer wenn auf "," (Komma dann Leerzeichen) das Wort "LLC" oder "Inc" folgt (dann nichts tun)?
Ich will es verändern:
- "TEXAS ENERGY MUTUAL, LLC, BOBBY GILLIAM, STEVE PEREIRA und ANDY STITT"
- "Grape, LLC, Andrea Gray, Jack Smith"
- "Stephen Winters, Apple, Birne, Inc, Sarah Smith"
Dazu:
- "TEXAS ENERGY MUTUAL, LLC_BOBBY GILLIAM_STEVE PEREIRA_ANDY STITT"
- "Grape, LLC_Andrea Gray_Jack Smith"
- "Stephen Winters_Apple_pear, Inc_Sarah Smith"
Ich dachte, es würde mit einer Variation des folgenden Codes beginnen, aber ich kann die Ausnahmebedingungen nicht herausfinden.
df ['Column_Name'] = df ['Column_Name']. str.replace (',', '_') Prost!
Antworten
1 ernest_k Oct 25 2020 at 05:35
Sie können einen regulären Ausdruck durch einen negativen Lookahead ersetzen :
#no idea why Inc|LLC or LLC|Inc will skip the first
df['Column_Name'].str.replace(', (?!=|Inc|LLC)', '_')
Ausgabe:
0 TEXAS ENERGY MUTUAL, LLC_BOBBY GILLIAM_STEVE P...
1 Grape, LLC_Andrea Gray_Jack Smith
2 Stephen Winters_Apple_pear, Inc_Sarah Smith
Name: ColumnName, dtype: object
1 chai Oct 25 2020 at 05:34
Verwenden Python Regex Modul re für mit dem Muster , (?!Inc|LLC)alle Vorkommen zu finden , , ohne auf IncoderLLC
import re
strings = ["Banana, orange", "Grape, LLC", "Apple, pear, Inc"]
[re.sub(", (?!Inc|LLC)",'_',string) for string in strings]
#['Banana_orange', 'Grape, LLC', 'Apple_pear, Inc']
DozParp Oct 25 2020 at 05:52
der einfache Weg:
def replace(str):
x = str.split(', ')
buf = x[0]
for i in range(1, len(x)):
if x[i].startswith('LLC'):
buf += ', ' + x[i]
elif x[i].startswith('Inc'):
buf += ', ' + x[i]
else:
buf += '_' + x[i]
return buf
und dann versuchen replace('a, b, LLC, d')